docling-parse

docling-parse

Extract text with coordinates from PDFs

Description

PDF parsing and layout analysis need each character and line's exact position on the page. docling-parse, a core Docling component, extracts text with coordinates from programmatic PDFs.

Implemented in C++ for speed and accuracy, with Python bindings.

Features



Coordinates:Exact positions.

Fast:C++.

Python:Easy use.