
Description
Training multi-turn search agents with only a final right-or-wrong reward hides which steps helped. IGPO, an ICLR 2026 work, uses the information gain of each turn as the reward signal, a simple, effective way to optimize search agents.
The repo includes the paper implementation.
Information gain:Feedback every turn.
Multi-turn search:Better retrieval agents.
Simple:Easy to implement.
The repo includes the paper implementation.
Features
Information gain:Feedback every turn.
Multi-turn search:Better retrieval agents.
Simple:Easy to implement.

