Giskard

Giskard

Open source evaluation and testing for LLM agents

Description

Whether a deployed AI agent hallucinates, shows bias or retrieves accurately cannot be covered by manual spot checks. Giskard is an open source evaluation and testing library for LLM agents that generates tests, runs evals and surfaces issues.

v3 is a fresh modular, lightweight, async-first rewrite for checking agents before release and in CI.

Features



Test generation:Automatic cases.

Evaluation:Hallucination, bias and accuracy.

RAG:Retrieval quality.

Async-first:v3 rewrite.