PDF Document Layout Analysis

PDF Document Layout Analysis

Docker service for PDF layout analysis

Description

Extracting body text from PDFs mixes up headings, tables, headers and footnotes. This project offers a Docker-powered PDF document layout analysis service that identifies each part of a page.

It segments pages into titles, text, tables, images and formulas with reading order, plus TOC and table extraction.

Features



Segmentation:Page regions.

Classification:Titles, tables and images.

Reading order:Correct flow.

Docker:Easy deploy.