
PDF Document Layout Analysis
Docker service for PDF layout analysis
Description
Extracting body text from PDFs mixes up headings, tables, headers and footnotes. This project offers a Docker-powered PDF document layout analysis service that identifies each part of a page.
It segments pages into titles, text, tables, images and formulas with reading order, plus TOC and table extraction.
Segmentation:Page regions.
Classification:Titles, tables and images.
Reading order:Correct flow.
Docker:Easy deploy.
It segments pages into titles, text, tables, images and formulas with reading order, plus TOC and table extraction.
Features
Segmentation:Page regions.
Classification:Titles, tables and images.
Reading order:Correct flow.
Docker:Easy deploy.

