Ollama-OCR

Ollama-OCR

OCR with vision models through Ollama

Description

Traditional OCR fails on complex layouts, handwriting and tables, and cloud OCR needs uploads. Ollama-OCR is an OCR tool extracting text from images and PDFs with vision language models through Ollama, all local.

It supports LLaVA, Llama 3.2 Vision and more, outputting Markdown, JSON or text as a Python package and Streamlit app.

Features



Vision models:Several.

Images and PDF:Batch.

Outputs:Markdown and JSON.

Local:Private.