Skip to main content
Cover the situations. Run them in a sandbox that can fail. Grade with a checked judge. The one-page PDF is at withwhile.com/while-simulation-engine.pdf.

Eight steps

1

Axes

What varies: tool, policy clause, world state, fault, persona, history. A situation is a point in that space.
2

Cover

Every pair of axis values appears together at least once (a pairwise covering array). Most failures are two-factor interactions [6].
3

Search

Five situation writers fill the grid. Each batch, the budget follows the writers that found new behavior [7]. Stops when nothing new turns up.
4

World

Tools answer from schema-shaped state. Same seed, same reply. Unknown id: not found. An argument copied from the schema instead of the customer: refused.
5

Rollout

N tasks × n phrasings × k samples. Every row keeps the policy version, temperature and per-token logprobs.
6

Grade

Conduct rules first, then the judge. The judge is scored against gold labels before its grades count.
7

Cut

SFT rows (reward=1, loss on agent turns). DPO pairs with margin. GRPO groups in the 20 to 80 percent band [4, 5]. Or a reward-model set.
8

Delta

Re-run the held-out tasks. Paired difference per task, bootstrap interval, permutation p [1].

Measurement

The held-out tasks are the only number that counts. Everything else on this page exists to make that number mean something.

Questions

Importance sampling? No. We are covering the failure space, not estimating production. Each row carries logprobs, policy version and temperature, so an off-policy trainer can form the ratio itself [10, 11]. SFT or RL? Both, from the same graded rows. Reward=1 rows for SFT, pairs for DPO, groups for GRPO, everything for a reward model. The judge is another LLM. Yes. So it is measured against gold labels, probed with known hacks, versioned by rubric hash, and drawn from a different model family than the policy [8]. How do you know training helped? Paired before and after on held-out tasks with a bootstrap interval. Markers such as argument grounding catch what pass@1 hides [1].

Code

The same eight steps with the exact estimators and knob names are on the engine on one page in the guides.

References

  1. Lambert, N. Reinforcement Learning from Human Feedback. 2025.
  2. Chen, M. et al. Evaluating Large Language Models Trained on Code. 2021.
  3. Yao, S. et al. τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. 2024.
  4. Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. 2024.
  5. Yu, Q. et al. DAPO: An Open-Source LLM Reinforcement Learning System at Scale. 2025.
  6. Kuhn, D. R., Wallace, D. R., Gallo, A. M. Software Fault Interactions and Implications for Software Testing. IEEE TSE 30(6), 2004.
  7. Lehman, J., Stanley, K. O. Abandoning Objectives: Evolution Through the Search for Novelty Alone. Evolutionary Computation 19(2), 2011.
  8. Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023.
  9. Gao, L., Schulman, J., Hilton, J. Scaling Laws for Reward Model Overoptimization. ICML 2023.
  10. Schulman, J. et al. Proximal Policy Optimization Algorithms. 2017.
  11. Noukhovitch, M. et al. Asynchronous RLHF: Faster and More Efficient Off-Policy RL for Language Models. ICLR 2025.