gImageReader

gImageReader

Graphical front-end for tesseract-ocr

Description

Want to extract text from scans, PDFs and images without wrestling with tesseract on the command line? gImageReader wraps it in an easy graphical interface—import images or PDFs, box areas to recognize, with multilingual support, spell-check, and searchable hOCR/PDF output. Fine-tune paragraphs and word blocks against the recognized result. Free and open source, a handy tool for batch OCR.

Features



Graphical OCR: an easy GUI for tesseract-ocr.

Many sources: supports images, PDFs and scanner sources.

Area recognition: box areas and fine-tune paragraphs/word blocks.

Searchable output: output hOCR, searchable PDF or plain text.