Camelot

Camelot

Extract tables from PDFs in Python

Description

Tables in financial reports and statistics PDFs come out garbled when copied, and retyping is slow and error-prone. Camelot is a Python library to extract tabular data from PDFs, turning tables into DataFrames in a few lines.

Two parsing modes handle ruled and unruled tables, exporting CSV, Excel, JSON and more.

Features



Extraction:PDF to data.

Two modes:Ruled and unruled.

DataFrames:Ready to analyze.

Export:CSV, Excel and JSON.