pdfplumber

pdfplumber

Extract text and tables from PDFs precisely

Description

Copying tables from PDF reports scrambles rows, and ordinary parsers lose positions. pdfplumber is a Python library that inspects every character, rectangle and line in a PDF to extract text and tables easily.

It gives precise coordinates, visual debugging, cropping and custom table rules.

Features



Tables:Accurate rows.

Characters:Positions and fonts.

Debugging:Visual.

Cropping:By region.