IGPO

IGPO

Information gain policy optimization for search agents

Description

Training multi-turn search agents with only a final right-or-wrong reward hides which steps helped. IGPO, an ICLR 2026 work, uses the information gain of each turn as the reward signal, a simple, effective way to optimize search agents.

The repo includes the paper implementation.

Features



Information gain:Feedback every turn.

Multi-turn search:Better retrieval agents.

Simple:Easy to implement.