Rex-Omni

Rex-Omni

Detect anything via next point prediction

Description

Classic detectors only know the classes they were trained on, so finding something new means labeling and retraining. Rex-Omni is a 3B multimodal LLM that turns object detection and other visual perception tasks into next-token prediction, so naming a category is enough to box it.

It also handles pointing, keypoints and OCR, with a Gradio demo, an AWQ quantized build and fine-tuning code, published at CVPR 2026.

Features



Open-vocabulary detection:Find objects by text.

Unified tasks:Detection, pointing, keypoints and OCR.

Quantized:AWQ halves storage.

Fine-tuning:SFT and GRPO.