
Description
Training LLM search skills with RL means repeated real search engine calls with high API costs and uneven results. ZeroSearch from Alibaba's Tongyi Lab incentivizes search capability with simulated search instead of a real engine, slashing training costs.
The repo includes the paper implementation, training code and models for retrieval and search agent research.
Simulated search:No real search API.
Cheap training:Big cost savings.
RL:Rewards retrieval skill.
Open models:With training code.
The repo includes the paper implementation, training code and models for retrieval and search agent research.
Features
Simulated search:No real search API.
Cheap training:Big cost savings.
RL:Rewards retrieval skill.
Open models:With training code.
Tags:llm

