
Description
LLMs make things up when their knowledge runs out, and fixed retrieve-then-answer is inflexible. Search-R1 is an RL framework that trains LLMs to call search engines on their own while reasoning.
Built on veRL for efficiency and scale, it teaches models when and what to search and how to use the results.
Interleaved:Reason and search.
RL:Learns to retrieve.
Efficient:Built on veRL.
Built on veRL for efficiency and scale, it teaches models when and what to search and how to use the results.
Features
Interleaved:Reason and search.
RL:Learns to retrieve.
Efficient:Built on veRL.

