wai.Harness treats the harness the way the rest of the library treats
weights: one object, a fingerprint that is its version, rows that say which
one produced them, and a report that says whether changing it did anything.
One object
The prompted loop the SDK plays itself: a model, the instructions, the tools. The fingerprint hashes what a comparison has to disclose [1] (model, instructions, tool names, and theDisclosure fields below), so a prompt
edit is a new version without anyone naming it [4].
Disclosure carries the rest of the setup: the context files the harness
loads, the turn cap, how it compacts a long context, retries, subagents,
sampling. Set a field and the hash changes; leave the label off and the
version is h- plus the hash.
prompt@model and the Runs page groups the dots by prompt
and by model as two axes.
A coding agent as the harness
Claude Code, Codex and pi each have a non-interactive mode that streams JSON events. A preset builds the command line, runs one subprocess per task, and normalizes the stream to the same{steps, final_text} every
adapter returns. The disclosure lists the context files present in cwd
(CLAUDE.md, AGENTS.md, .pi/SYSTEM.md), so the same command in a
directory with a different CLAUDE.md is a different harness.
claude -p and codex exec and pi --mode json each need their own login and the CLI on PATH.
Harness.command([...]): the token "{prompt}" in
the command line is replaced by the task text (or the text goes to stdin),
and parse= turns stdout into a trajectory. An agent you already have as a
Python callable is Harness(agent=fn, ...), given a label and a model name
so its rows can be compared.
Run it, and every row says so
simulate(harness) takes the tools, the system prompt and the turn cap
from the harness when the call does not name them, plays a prompted
harness through the engine (so the mock world and the scheduled faults
apply) or a command harness as it is, and stamps every row with
harness = {label, hash, model, kind}.
Which lever moved the score
Run every harness in a set on every model in a set over the same frozen tasks (tasks= the first run, so the asks match), grade them with one
judge, and hand all the rows to wai.harness.attribute. It reads the
harness x model grid off the rows and decomposes the spread of the cell
means into a harness part, a model part and their interaction, with a
bootstrap interval over tasks on each share [6]. It also says whether the
leading model changes from one harness to another, the ranking reversal of
[1]. The verdict is in words: the harness moved the score more than the
model did, the model did, or the difference could be chance.
The rows below are built by hand so the page runs offline; yours come out
of simulate with the stamp already on them.
Train under several harnesses
A policy trained under one fixed harness collapses when the tool environment shifts; one trained across harnesses holds up out of distribution [3].export_environment(harnesses=[...]) writes the harnesses
into the environment’s spec.json (label, hash, instructions, tool
schemas, disclosure), and the trainer’s environment draws one per task
from a hash of the task id and a seed, so a re-run draws the same map. The
rollout runs with that harness’s instructions as its system prompt and its
tool schemas as its tool set, and its state carries harness = {label, hash}, so a trace says which one it ran under. The package README lists
them under “Harnesses”. load_environment(harness_mix=) takes "uniform"
or one weight per harness.
On the platform
harness.pin() is the platform record with the same label and hash, and
track(harness=), tracked.run(harness=) and HarnessSweep accept the
runnable object directly. The wire does not change: the disclosure folds
into the hash and is not sent, and a platform Harness with no disclosure
keeps the hash it always had. This needs WHILEAI_API_KEY.
What is tested, and what is not yet
- The Claude Code preset ran live on 2026-09-21 (
haiku, two turns, one allowed tool): the stream parsed and the reply came back in five seconds. Thecodexandpiparsers are written from each CLI’s own documentation of its JSON stream and checked against recorded shapes intests/api/test_harness.py; nobody has run them live yet. If you do, open an issue with the first three lines of the stream. - The harness draw in an exported environment is tested offline through
verifiers’ own rollout state (
tests/api/test_environment.py): the same task draws the same harness twice, both harnesses appear across forty tasks, and the rollout’s system prompt and tools are that harness’s. No prime-rl run has trained on a two-harness spec yet. - Not here yet, tracked in #712: a harness from the Prime Intellect Environments Hub by id [5], and the Meta-Harness outer loop as a recipe [2].
References
- Zhang, Wang, Ge, Xu, Hamm, Reddy. Stop Comparing LLM Agents Without Disclosing the Harness. 2026. arXiv:2605.23950.
- Lee, Nair, Zhang, Lee, Khattab, Finn. Meta-Harness: End-to-End Optimization of Model Harnesses. 2026. arXiv:2603.28052.
- Kim, Choi, Lee, Jun, Kim, Park. The Interplay of Harness Design and Post-Training in LLM Agents. 2026. arXiv:2606.25447.
- Lambert. Reinforcement Learning from Human Feedback, chapter Evaluation. 2025. rlhfbook.com.
- Prime Intellect. verifiers v1: Decomposing Tasksets and Harnesses for Agentic RL & Evaluations. 2026. primeintellect.ai/blog/verifiers-v1.
- Miller. Adding Error Bars to Evals. 2024. arXiv:2411.00640.