OpenAI Evals

OpenAI Evals

Framework for evaluating LLMs and LLM systems

Description

After changing a prompt or model, judging whether things improved by feel is unreliable. OpenAI Evals is a framework for evaluating LLMs and LLM systems and an open registry of benchmarks.

It supports custom evals, model-graded scoring and batch runs to turn your use cases into repeatable tests.

Features



Framework:Custom evals.

Registry:Open benchmarks.

Model-graded:Auto scoring.

Repeatable:Regression checks.