Notes Sep 29, 2026 1 min read

Why the evaluation set came before the prompt

pythonevaluationcillm

RepositoryRun locallyEvaluation setDecision records

Context

It's tempting to start an LLM feature by tuning the prompt until the demo looks right. That tells you the demo works. It doesn't tell you which failures the system catches, or whether next week's change quietly stopped catching one.

Decision

The cases for AI Compliance Agent are nine small, versioned fixtures in evals/cases/v1.json. Each has a document, the policies available, a fixed model output, and the errors the system must report. evals/run.py runs them through the real retrieval and validation code and prints expected versus actual results. CI runs that script on pull requests and pushes to main. The runner uses a fixed output provider, so it needs no API key and gives the same answer every time.

What I rejected

Grading these checks with another model. Whether a quote is in the policy, and whether a date or duration in the rationale is in the cited text, has an exact answer, so evals/run.py compares expected and actual errors in code.

Failure scenario

The trivially-short-quote case expects citation_quote_too_short. CI runs it. A change that lets a quote shorter than 20 characters pass fails that case.

Limits

The nine cases are fixtures in the repository. They show the checks behave as specified. They use a fixed model output, so they are not a measure of model accuracy on real documents.

Evidence

Inspect: evals/cases/v1.json · evals/run.py · CI