
Description
Extracting fields from invoices and contracts with classic OCR yields raw text, and cloud services are off-limits for sensitive files. docext from Nanonets is an on-premises document extraction and benchmarking toolkit that skips traditional OCR and uses vision language models.
It converts documents to Markdown, extracts tables and fields, and comes with the 3B Nanonets-OCR-s model that understands images, signatures and watermarks.
On-premises:Sensitive docs stay inside.
Field extraction:From unstructured documents.
Tables and Markdown:Structured output.
Dedicated model:Nanonets-OCR-s 3B.
It converts documents to Markdown, extracts tables and fields, and comes with the 3B Nanonets-OCR-s model that understands images, signatures and watermarks.
Features
On-premises:Sensitive docs stay inside.
Field extraction:From unstructured documents.
Tables and Markdown:Structured output.
Dedicated model:Nanonets-OCR-s 3B.
