
Description
Classic detectors only know the classes they were trained on, so finding something new means labeling and retraining. Rex-Omni is a 3B multimodal LLM that turns object detection and other visual perception tasks into next-token prediction, so naming a category is enough to box it.
It also handles pointing, keypoints and OCR, with a Gradio demo, an AWQ quantized build and fine-tuning code, published at CVPR 2026.
Open-vocabulary detection:Find objects by text.
Unified tasks:Detection, pointing, keypoints and OCR.
Quantized:AWQ halves storage.
Fine-tuning:SFT and GRPO.
It also handles pointing, keypoints and OCR, with a Gradio demo, an AWQ quantized build and fine-tuning code, published at CVPR 2026.
Features
Open-vocabulary detection:Find objects by text.
Unified tasks:Detection, pointing, keypoints and OCR.
Quantized:AWQ halves storage.
Fine-tuning:SFT and GRPO.

