Terminal-Bench

Terminal-Bench

Benchmark for frontier agent work in the terminal

Description

Everyone claims the best coding agent, but which actually finishes work in a terminal? Terminal-Bench is a benchmark measuring AI agents on hands-on terminal tasks, diverse, difficult and continuously evolving.

Frontier agent builders use it to track progress and compare, with releases on Harbor Hub.

Features



Real tasks:In the terminal.

Hard:Frontier-level.

Evolving:Tagged releases.

Comparison:Leaderboard.