Eight steps
1
Axes
What varies: tool, policy clause, world state, fault, persona, history.
A situation is a point in that space.
2
Cover
Every pair of axis values appears together at least once (a pairwise
covering array). Most failures are two-factor interactions [6].
3
Search
Five situation writers fill the grid. Each batch, the budget follows
the writers that found new behavior [7]. Stops when nothing new turns
up.
4
World
Tools answer from schema-shaped state. Same seed, same reply. Unknown
id: not found. An argument copied from the schema instead of the
customer: refused.
5
Rollout
N tasks × n phrasings × k samples. Every row keeps the policy version,
temperature and per-token logprobs.
6
Grade
Conduct rules first, then the judge. The judge is scored against gold
labels before its grades count.
7
Cut
SFT rows (reward=1, loss on agent turns). DPO pairs with margin. GRPO
groups in the 20 to 80 percent band [4, 5]. Or a reward-model set.
8
Delta
Re-run the held-out tasks. Paired difference per task, bootstrap
interval, permutation p [1].
Measurement
The held-out tasks are the only number that counts. Everything else on
this page exists to make that number mean something.
Questions
Importance sampling? No. We are covering the failure space, not estimating production. Each row carries logprobs, policy version and temperature, so an off-policy trainer can form the ratio itself [10, 11]. SFT or RL? Both, from the same graded rows. Reward=1 rows for SFT, pairs for DPO, groups for GRPO, everything for a reward model. The judge is another LLM. Yes. So it is measured against gold labels, probed with known hacks, versioned by rubric hash, and drawn from a different model family than the policy [8]. How do you know training helped? Paired before and after on held-out tasks with a bootstrap interval. Markers such as argument grounding catch what pass@1 hides [1].Code
The same eight steps with the exact estimators and knob names are on
the engine on one page in the guides.
References
- Lambert, N. Reinforcement Learning from Human Feedback. 2025.
- Chen, M. et al. Evaluating Large Language Models Trained on Code. 2021.
- Yao, S. et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024.
- Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. 2024.
- Yu, Q. et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025.
- Kuhn, D. R., Wallace, D. R., Gallo, A. M. Software Fault Interactions and Implications for Software Testing. IEEE TSE 30(6), 2004.
- Lehman, J., Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation 19(2), 2011.
- Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
- Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023.
- Schulman, J. et al. Proximal Policy Optimization Algorithms. 2017.
- Noukhovitch, M. et al. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models. ICLR 2025.