DeepEyesV2

DeepEyesV2

Toward an agentic multimodal model

Description

Multimodal models answering about images guess at unclear details and never look things up. DeepEyesV2 is an agentic multimodal model that thinks with images, cropping and zooming, running code and searching the web while reasoning.

Trained with cold start and RL, with checkpoints available, it excels at visual tasks needing close looks and verification.

Features



Thinking with images:Crop and zoom.

Tools:Code and search.

RL:Learns when to use tools.