
Description
Everyone claims the best coding agent, but which actually finishes work in a terminal? Terminal-Bench is a benchmark measuring AI agents on hands-on terminal tasks, diverse, difficult and continuously evolving.
Frontier agent builders use it to track progress and compare, with releases on Harbor Hub.
Real tasks:In the terminal.
Hard:Frontier-level.
Evolving:Tagged releases.
Comparison:Leaderboard.
Frontier agent builders use it to track progress and compare, with releases on Harbor Hub.
Features
Real tasks:In the terminal.
Hard:Frontier-level.
Evolving:Tagged releases.
Comparison:Leaderboard.
