
Description
Agents pick wrong tools or mangle parameters, and it's unclear which model is more reliable. ToolCall-15 is BenchLocal's benchmark pack for LLM tool calling.
It tests selection, parameter precision, multi-step chains, restraint and error recovery, compared side by side in BenchLocal.
Selection:Right tool.
Parameters:Precise.
Chains:Multi-step.
Recovery:From errors.
It tests selection, parameter precision, multi-step chains, restraint and error recovery, compared side by side in BenchLocal.
Features
Selection:Right tool.
Parameters:Precise.
Chains:Multi-step.
Recovery:From errors.
