
Description
A multimodal model that reads images, documents and video usually means pricey closed APIs and sending data away. InternVL from Shanghai AI Lab is an open source multimodal chat model family approaching GPT-4o.
It spans sizes from 1B to tens of billions of parameters, excels at OCR, charts and video Q&A, and can be deployed and fine-tuned locally.
Vision-language:Image Q&A.
OCR:Tables and charts.
Video:Long-video Q&A.
Sizes:Edge to server.
It spans sizes from 1B to tens of billions of parameters, excels at OCR, charts and video Q&A, and can be deployed and fine-tuned locally.
Features
Vision-language:Image Q&A.
OCR:Tables and charts.
Video:Long-video Q&A.
Sizes:Edge to server.
