
Description
Whether a deployed AI agent hallucinates, shows bias or retrieves accurately cannot be covered by manual spot checks. Giskard is an open source evaluation and testing library for LLM agents that generates tests, runs evals and surfaces issues.
v3 is a fresh modular, lightweight, async-first rewrite for checking agents before release and in CI.
Test generation:Automatic cases.
Evaluation:Hallucination, bias and accuracy.
RAG:Retrieval quality.
Async-first:v3 rewrite.
v3 is a fresh modular, lightweight, async-first rewrite for checking agents before release and in CI.
Features
Test generation:Automatic cases.
Evaluation:Hallucination, bias and accuracy.
RAG:Retrieval quality.
Async-first:v3 rewrite.

