
Description
Traditional OCR fails on complex layouts, handwriting and tables, and cloud OCR needs uploads. Ollama-OCR is an OCR tool extracting text from images and PDFs with vision language models through Ollama, all local.
It supports LLaVA, Llama 3.2 Vision and more, outputting Markdown, JSON or text as a Python package and Streamlit app.
Vision models:Several.
Images and PDF:Batch.
Outputs:Markdown and JSON.
Local:Private.
It supports LLaVA, Llama 3.2 Vision and more, outputting Markdown, JSON or text as a Python package and Streamlit app.
Features
Vision models:Several.
Images and PDF:Batch.
Outputs:Markdown and JSON.
Local:Private.

