Skip to main content
Source Code: src/gaia/eval/

What Is Agent Eval?

The Agent Eval framework validates your GAIA agent’s quality by running realistic, multi-turn conversations against the live Agent UI. It uses Claude Code as both user simulator and judge — driving conversations through MCP, scoring every response across 7 dimensions, and producing a machine-readable scorecard. Unlike unit tests that check individual functions, Agent Eval tests the full system end-to-end: RAG indexing, tool dispatch, context retention, hallucination resistance, and personality compliance — through the same interface real users interact with. What you get:
  • Automated multi-turn conversation testing with persona-driven user messages
  • 7-dimension scoring rubric (correctness, tool selection, context retention, completeness, efficiency, personality, error recovery)
  • Deterministic pass/fail with weighted scoring
  • Regression detection via baseline comparison
  • Auto-fix mode that invokes Claude Code to repair failures

Prerequisites

1

Install eval dependencies

2

Set your Anthropic API key

The benchmark uses Claude as the judge model:
3

Install Claude Code CLI

The runner invokes scenarios via claude -p subprocess. Verify it’s installed:
If not installed, see Claude Code installation.
4

Start the LLM backend

Lemonade Server provides the local LLM and embeddings:
5

Start the Agent UI backend

The Agent UI backend must be running on http://localhost:4200 (default).

Quick Start

Run your first eval in under 5 minutes:

1. Run a Single Scenario

Start with a simple RAG factual lookup test:
This will:
  1. Create a new Agent UI session via MCP
  2. Index the acme_q3_report.md document
  3. Ask about Q3 revenue (simulating a power_user persona)
  4. Judge the response against the ground truth ($14.2 million)
  5. Output a scored result

2. Run a Category

Test all RAG quality scenarios:

3. Run the Full Benchmark

Run all scenarios across every category (91 scenarios in 12 categories at the time of writing):

4. Architecture Audit (Free)

Check for structural limitations without making any LLM calls:

Reading the Scorecard

After each run, results are written to eval/results/<run_id>/:

Understanding Pass / Fail

A scenario passes when both conditions are met:
  • Overall score is 6.0 or higher (out of 10)
  • No turn has a correctness score below 4
A scenario fails if either:
  • Overall score is below 6.0, OR
  • Any single turn has correctness below 4 (hard fail on hallucination or wrong answer)

The 7 Scoring Dimensions

Each turn is scored across 7 dimensions. The overall score is a weighted sum:

Status Codes

Only PASS, FAIL, and BLOCKED_BY_ARCHITECTURE count toward the average score and judged pass rate. Infrastructure statuses are excluded from quality metrics.

Sample Output


Scenario Categories

The benchmark includes 91 scenarios across 12 categories (counts as of this writing — the runner discovers scenarios dynamically, so totals grow as new YAML is added):

Common Workflows

Regression Testing

Save a baseline after a known-good run, then compare future runs:
The comparison shows per-scenario deltas, regressions (PASS to FAIL), and improvements (FAIL to PASS).

Auto-Fix Mode

Let Claude Code automatically diagnose and repair failures:
Fix mode runs in a loop:
  1. Evaluate all scenarios
  2. Diagnose failures and patch source code
  3. Re-run only the failed scenarios
  4. Compare results — stop when --target-pass-rate is reached
Fix mode patches src/gaia/ source files directly. Always review diffs before committing. Run python util/lint.py --all --fix after fix iterations.

Capturing Real Sessions

Convert a live Agent UI conversation into a replayable scenario:
This reads the session from the Agent UI database, extracts turns and indexed documents, and writes a scenario YAML to eval/scenarios/captured/. You must then edit the file to add proper ground_truth and success_criteria fields.

Filtering by Tags

Run only scenarios with specific tags:

Output Formats


Cost Control

Typical costs:
  • Single scenario: 0.020.02 -- 0.10
  • Full benchmark (all scenarios): 1.001.00 -- 5.00 (scenarios whose corpus files aren’t on disk are skipped)
  • Architecture audit: $0.00 (no LLM calls)
Start with --audit-only (free) to understand expected failures, then run individual categories to control costs during development.

Next Steps

Scenario Authoring

Write custom scenarios with YAML and ground truth

CI/CD Integration

Run Agent Eval in GitHub Actions with regression detection

CLI Reference

Complete flag reference for gaia eval agent

Agent Eval Benchmark

Deep-dive into architecture, scoring pipeline, and fix mode internals