recipes/02-measure/eval-your-agent
(offline, no key, seconds).
Names first, because they cost testers ten minutes: zp, ZeroProof and
While are the same product. The package is whileai
(pip install whileai; pip install zeroproof still works as a shim),
the import is whileai.simulations, API keys start with zp_, and the
docs live at docs.withwhile.com.
1. Install and sign in
simulator=False. On a
machine that had the old package, ~/.zeroproof/credentials.json is
still picked up as long as ~/.whileai holds no credentials; set
WHILEAI_HOME=/some/fresh/dir to isolate a new account from it.
The old names still work, so you may already be signed in.
ZEROPROOF_API_KEY is read whenever WHILEAI_API_KEY is unset (every
ZEROPROOF_* variable is), a saved ~/.zeroproof/credentials.json
counts as a login, and pip install zeroproof installs whileai. If
whileai status names a key you never set here, that is where it came
from.signup key is a trial: 25,000 input
and 50,000 output tokens a day, which is about twelve hosted situations
of a four-tool agent. One real run spends that, and the run then stops
with Hosted model daily quota exceeded. Two ways around it:
simulate(..., simulator=False) writes the situations offline with no
quota and no network, which is how every recipe here runs; and signing
in once at the While site (the link whileai status prints) lifts the
daily limit. whileai status prints the same two facts while the key is
on the trial, and a run that would spend the trial on the hosted writer
says them once before it starts, in the log and in data.warnings.
2. Wrap your agent
The engine calls your function once per rollout with the ask the writer produced, and wants back the tool calls it made and what it said:- It runs its own real tools. The engine’s mock world and scheduled faults apply to model-backed agents; your callable answers its own calls. That is what you want for an eval of the thing you ship.
- It is played single-turn. One message in, one trajectory out. The
multi-turn user model (
avg_turns,max_turns) does not apply. - The writer does not know your ids. It reads the tool descriptions
and the system prompt. If your world has order numbers or account
names, put them in the tool description (“Orders on file: A1001,
A1002, …”) or in
seeds=, or the writer invents ids, every rollout is “not found”, and the run is hollow.
{"type": "function", "function": {"name", "description", "parameters"}});
a bare {"name", "description", "parameters"} dict works too, and so does
the Anthropic shape {"name", "description", "input_schema"}
(input_schema is read as parameters). If your bot records calls
through a shared global, wrap the recorder in a threading.local:
concurrency defaults to 32, so 32 threads call your function at once
and one shared list interleaves calls from different rollouts into each
other’s rows.
Or run whileai init-evals in the project and edit the files it writes.
It reads your Python with ast, never imports it, picks the tool list,
the system prompt and the callable that answers a message, and writes
five files: evals/agent.py (this wrapper, with your tools converted to
OpenAI shape and your tool runner wrapped in the thread-local recorder),
evals/judge.py, evals/run.py, evals/test_judge.py and
evals/README.md, wired to each other. It prints what it picked, so a
wrong guess is one flag away: --agent module:callable,
--tools module:NAME, --system-prompt module:NAME. When it finds
nothing the files are still written, with every place that needs your
code marked TODO. It ends with the three commands to run next:
python evals/run.py --gap, python evals/run.py, and
pytest evals/test_judge.py.
3. Write the judge as a program
The judge reads the trajectory, not the prose. A polite reply that issued a refund it should not have scores 0; a blunt one that followed the policy scores 1. The policy lives in the judge, once:reward in [0, 1], a reason string, and optional
markers (name to 0/1). Name markers so 1.0 is always the good outcome
(refund_only_when_allowed, not refunded_wrongly); the marker table
reads as one column then. A marker that does not apply to a row is
None, so its rate counts only the rows it measured. A verifier
(wai.verify.*) is a judge too, when the answer is checkable: each one
is a callable that takes a row and returns this same dict.
Any other key you return is kept under row["judge_meta"], not on the row:
a judge that returns failures reads back as row["judge_meta"]["failures"].
3b. Find what your tests miss
The suite you have sends a set of asks. Which parts of the policy do they never reach?! note per gap:
asks is a list of prompt strings, a list of rows with a prompt key, or
a path to a .py or .jsonl file holding either. From a .py file the
asks are the string literals that look like asks (passed to a call or in a
list, over fifteen characters, with a space): a heuristic, so read
report["asks"] before trusting the counts.
The axes are the ones simulate covers, so the report is in the engine’s
own words: untested_rules are the policy clauses no ask reaches,
untested_tools the tools no ask names, single_shot says every ask runs
once (one rollout cannot tell a flake from a failure), and notes names
the fix for each. world_state and tool_condition are not readable from
an ask at all, which is the honest reason a hand-written suite misses
fault handling: a prompt never says the record is missing or the tool
timed out.
Rules are matched on the words an ask shares with the clause, so a branch
that only the fixture data selects (an amount, a date) reads as untested
even when an ask lands on it. Pass rows= from a graded run to check the
world side: a rule whose every row ended in the same tool fault is one the
asks reach but the fixtures never let happen, and the fix is a fixture
case, not another ask.
preflight(tools, system_prompt)["rules"] is the same rule axis on its
own, which is the list of policy branches the engine extracted from your
prompt.
4. Run it
pass@1is how often the agent does the job.pass^kis how often it did on every one ofktries: for anything that moves money, that is the number.pass@kminuspass@1is headroom for training.repeat_policy="fixed"asks for all repeats up front. Themode="rl"default,"successive", stops early on unanimous asks, which is the right economy for training data and the wrong one for an eval.- Slice by category: tag each row (
row["category"] = classify(prompt)) and callwai.pass_aton each slice. The recipe prints that table. simulator=Falseis the offline template writer (no key, no network);simulator="hosted"is the default, the hosted writer, and means the same as leaving the argument out.pass^kandpass@kprint asn/a(the fields areNone) belowrepeats=4(min_k), because four tries is the smallest draw those numbers mean anything on; the line ends withset repeats>=4 for pass^k and pass@k, which is also.note.situationscounts asks,budgetcounts rows. Keepbudget >= situations * repeatsor the run stops at the budget with the later situations never rolled out at all.- To measure a policy branch, pin the tool result. A rule like
“credits over 200, and the situation writer invents the
amount, so on an unforced run the branch is reached at random: a
marker that never fires, or reads 1.000 because it never had a chance
to fail. Pin it:
wai.local_model(..., result_shapes={"lookup_invoice": {"invoice_id": "INV-1000", "amount_usd": 900.0, "status": "open"}}). Numbers move by up to about a third per call (900.0lands in roughly 600 to 1200,90.0in 60 to 120), so pick a value whose whole range sits on one side of the threshold, and run the same pinnedtasks=once per side. Measured on a billing agent: 44 lookups over $200 and 0 under with the big shape, 48 under and 0 over with the small one, no leakage across 34 tasks x 4 rollouts. - A served model that scaled to zero takes two to three minutes to
answer its first request.
timeout=is 300 s by default so that first pass lands; if a call still times out,data.warningssays so and names the fix (raisetimeout=, or send one throwaway request first). A pass that comes back with fewer rows than the base arm is the dangerous case, since some tasks then sit at k=1 against the base’s k=4; readdata.warningsbeforepass_at. - A long run says where it is on the
whileai.simulationslogger, one line at most ten seconds and ten rollouts apart (12/64 rollouts, 3 situations written, 1m40s elapsed, ~5m left); runs under ten rows say nothing. Calllogging.basicConfig(level=logging.INFO)to see it; a hosted run can sit a minute before the first row, and silence is not a hang.
5. Gate CI on it
Two lanes. The slow one runs the agent (model calls, seconds to minutes) and exits non-zero under a floor:6. Check the judge
A judge is a claim until it is measured. Label a sample by hand, attach the labels as human, and ask:labels is a {key: 0/1} dict, a list of dicts, or a JSONL path. The
key names the row: rollout_id when the row has one, else
scenario_id#rollout_index (what a simulate row carries), else the
row’s prompt plus final_text.
judge_trust reads gold_reward (which attach_labels writes) and
reports agreement with its Wilson lower bound (the bottom of the 95%
interval), held-out halves, a length bias check and re-judge flips.
FAIL on eight labels means label more, not that the judge is wrong:
at perfect agreement the lower bound needs sixteen labels to clear 0.8,
and the report line says how many to label. Labels attached any other
way count as model-made and keep ok false unless you say
allow_model_gold=True.
7. Return shapes
The names in the print and the names on the object are not always the same word, and guessing costs a round trip. What each call hands back:
Two concepts here have two spellings each. These docs use the left one;
the right one is the same thing under another name, and the recipe uses
it in places.
PassAt fields, with the name each prints as:
Marker stats (
marker_summary(rows)["grounded"]): metric, mean,
ci95 (not ci), n_tasks, n_rows (not n), n_rows_at_1,
n_rows_at_0, degenerate, and one of note (no interval, too few
tasks: the bootstrap needs three) or warning (the marker never varied).
ci95 is None in both cases, and the sentence says which one you have.
judge_trust(rows): ok (measured and clean), agreement.agreement
with agreement.ci95, agreement.kappa and agreement.n, gold_kind
("human", "model", "unknown"), n_labeled, held_out_halves,
length_sensitivity, perturbation, probes, disagreements, and
warnings, where every line names its own fix.
8. Then
- Every failure row is a training example:
simulate(traces=scored.failures())aims the next round at what broke.evaluaterows are stampedlineage.source == "eval", and every selector report counts them aseval_sourcedand warns before you train on them. - Push the eval set with a purpose so it stays out of training:
scored.push("refund-evals", purpose="eval"). - Production traces are rows too:
wai.rows_from_otel(spans)reads OpenTelemetry spans, and the same judge and markers score them, so the number you got here is the number you watch after shipping.