import whileai.simulations as wai, then wai.name.
simulate
repeats (rollouts per ask, k), phrasings (wordings per
situation, n), repeat_policy ("fixed" gives every ask all k
repeats; mode="rl" defaults to "successive", which stops
early on unanimous asks), concurrency (parallel rollouts, 32),
simulator (the situation writer; False is the offline
template writer, no key, and "hosted" is the default written out,
the same as leaving it unset), user_model, backend (the agent’s
model when agent= is not a callable), seed, sampling,
max_turns / avg_turns (model-backed agents only; a callable
agent is played single-turn: one message in, one trajectory out),
fault_rate, temperature, timeout, logprobs. Each
is None unless you set it, and a misspelled keyword is a
TypeError, never silently ignored. seeds= is a list of
opening asks the writer keeps and varies; with a callable agent
whose world has real ids (order numbers, account names), put those
ids in the seeds or the tool descriptions, or the writer invents
ids and every rollout is “not found”.
Every seed is run. Each one becomes at least one situation:
situations is sized up to len(seeds) when you pass a smaller
number, and the search never spends a seed’s slot on an ask it wrote
itself. The one thing that can still drop a seed is budget,
which pays for len(seeds) * repeats rows before anything else;
when it cannot, the run says which seeds it dropped in warnings
and lists them in search["seeds_dropped"] before rolling out.
Seeds are asks to build a run around, not the eval set: to check
that a fixed list of asks all ran and how each scored, use
evaluate(eval_set=asks).
Input is an intent or an agent: system_prompt alone, tools
plus a prompt, or spec=. Search writes a grid of human requests
(ordinary, vague, complex, adversarial) and a spread of agent replies.
It spends budget rows and time_budget seconds on new coverage.
No default reward. Pass grade=True for the deterministic
conduct grade, or grade later with grade(). A callable
grader= is the only in-simulate score hook.
traces= (rows or a JSONL path of production traces) aims the
covering grid at the behaviors those traces show instead of the whole
space, and drops any generated row that near-copies a source trace,
so held-out traces stay out of training. Without it the grid comes
from the agent’s tools and policy alone (cold start). Traces
reproduce situations: the tools, faults and world states the agent
met. A failure mode that lives in how the reply is worded (an
unsupported claim, an estimate not labelled as one, two questions
where one was asked for) has no world-visible trigger, so traces
alone cannot aim at it; put a grader in the loop for those.
With grader= set, the grader’s verdict steers the search the way
a tool fault already does: a row the grader failed is re-rolled and
its ask is mutated into new ones, so the budget moves toward what
the grader catches, not only toward broken tools. A grader that
fails reply-form rules is exactly the signal the loop was missing
(#285: with a 12-rule grader, every rule with a tool-result trigger
was reproduced and every rule about the reply’s wording was not),
and without a grader there is no verdict to steer by, so there is no
switch to set: the grader is the switch. A graded failure is a reward
under 0.5 (a 0 from a 0/1 judge, a failed verifier, a rubric below
half); markers ride along on the row but do not aim on their own,
since their direction differs per marker. search["mutation_aims"]
counts the parents and the mutated rows per aim, world_fault and
graded_failure. To grade beside the loop and still steer by tool
faults alone, advanced={"mutate_graded_failures": False}.
hard_share= is the difficulty dial: the fraction of situations drawn
from the ambiguous, boundary and adversarial tiers (default 0.40), where
a base fails most often. search["tier_mix"] reports the share asked
for and the share drawn; dimensions={"stance": [...]} pins the axis
and keeps the other axes of the grid.
execute= is the caller’s world: a function (tool, arguments) -> result that answers every tool call for real, against their repo,
database, or service. Without it the mock world answers, which fits
record-shaped tools and not code. Scheduled faults still apply first.
whileai.simulations.generate.agents.current_rollout is a
thread-local set before each rollout with prompt, rollout_index
and seed, so execute can tell which run it is answering.
Three models can take part: the agent (agent= / backend=), the
situation writer (simulator=), and the simulated user
(user_model=, a backend spec; None means the writer’s model, the
agent’s own by default). Every row records all three next to
model_version: writer_model, user_model, and
judge_meta.model once graded. When the agent model also wrote the
situations or played the user, the run’s degraded list carries
same_model and warnings says which call separates them.
Training on a model’s own unfiltered output teaches it its own habits
(rlhf-book ch. 12), so the row says who wrote what.
patience= is how long the simulated person keeps answering the
agent’s questions. "normal" (the default) always tries to answer
the first question, and from the second question on may walk away
(35% on the second, 60% on each after that, drawn per thread so a
seeded run reproduces); at any question the person may also leave
when it asks for something they could not or would not know.
"short" walks away sooner (60% then 90%); "endless" never
walks away, which is what every run did before this knob: the
person answered every question until the depth cap, so no rubric
criterion about asking could fail. The odds are a default, not a
measurement: to ground them, fit a Kaplan-Meier hazard per question
index on source traces and set the levels from it. A row the person
left carries ended_by="user_left" and ends on the agent’s
question; search["ended_on_question"] is \{"share", "n", "user_left"\}: of n rows, the share that ended on a question and
how many of those the person left.
scaffold= is generation-only guidance appended to the system prompt
of the MODEL-BACKED teacher during rollout (and to the scene writer).
It never enters profile.policy, so exports and evals stay on the
plain policy; it is ignored for user-supplied callable agents. Measured
to help some agents and hurt others. Configure per agent, no default.
Every row says how it was sampled under sampling: temperature,
max_tokens and model, as the model backend resolved them
(rlhf-book ch. 16: a result is only comparable with its sampling
settings on record). A callable agent samples however it samples, so
its rows carry sampling: None unless you pass sampling={...},
which is recorded on every row as given.
Variation is three independent counts. Do not collapse them.
situations (N) is distinct worlds. requests_per_situation /
phrasings (n) is different human wordings of one world.
rollouts_per_request / repeats (k) is independent agent runs of
the same wording. Follow-ups branch on that run.
unique_situations=True keeps picking new worlds (n=1, k=1 unless you
set them). Silent aliases: n / phrasings for n, repeats /
rollouts_per_prompt for k, unique for unique_situations,
policy for system_prompt. Writer completions are
advanced["completions_per_request"]. Seed openers are
advanced["seed_prompts"].
tasks= re-runs a previous run’s task set instead of drawing a new
one: pass that run (SimulationData), its rows, or its JSONL path.
Every distinct prompt is rolled out again, repeat_count times,
on its own scenario_id and scenario_dimensions and under the
same faults and world state, and nothing else is generated.
repeat_count is k as this call resolves it — from this call’s
repeats / rollouts_per_request when given, otherwise from the
pinned run (the most rollouts any of its prompts has), never from
this call’s mode preset. So a base run made with repeats=4
and re-run as simulate(..., tasks=base) comes back at k=4 and
pass_at reports the same k on both sides; pass repeats= to
re-run at a different k on purpose::
base = wai.simulate(agent, tools=TOOLS, system_prompt=P,
mode=“rl”, repeats=4)
rerun = wai.simulate(agent, tools=TOOLS, system_prompt=EDITED,
tasks=base, mode=“rl”) # k=4, inherited
runs= replays the same task set that many times in one call and
stamps lineage.eval_run (0, 1, 2, …) on every row, which is
what delta_report needs before it will call a change real
(rlhf-book ch. 16 and appendix C: one evaluation is a draw, three
give a standard deviation). simulate(tasks=base, runs=3) is the
usual form; without tasks= the first run draws the task set (from
seeds= when given) and the rest replay it. Between runs nothing
changes but the agent’s own
sampling: same tasks, same faults, same world state, same seed, so a
deterministic agent gives identical runs and a zero re-run band. The
rows of every run come back in one SimulationData (output=
holds them all); search["eval_runs"] lists the rows and stop
reason per run, and eval_variance(data.rows()) splits by
eval_run on its own. budget is a per-run cap: runs=3, budget=100 returns up to 300 rows, and report()["budget_per_run"]
carries the cap under that name. Replayed rows keep the writer of the
run they replay on writer_model and say lineage.replayed_from_run,
so delta_report on two runs of one call sees one writer.
With situations=N the run stops once every one of the N
situations has all its rollouts (stopped_because= "situations_exhausted"), whatever budget still allows; a budget
above situations x phrasings x repeats is not spent.
A run
otherwise draws its tasks from the grid by seed and, above
concurrency: 1, by completion order, so a re-run shares only part
of its tasks with the first and compare_runs drops the rest;
pinning is how an A/B (a prompt edit, a model swap, another seed)
keeps every pair. The run stops when every pinned prompt has its
rollouts (stopped_because="tasks_done") or the budget is spent.
A seeded run is reproducible bit-for-bit at concurrency: 1,
apart from timing fields and per-invocation identity: with
grader= every row’s lineage.scoring_run_id names that one
scoring pass (a fresh id per call, like a timestamp), so two runs
differ on that key and on nothing else. Pass run_id= to
run_judge to pin it.
reproducible=True makes it so at any concurrency: each batch of
rollouts finishes, and every verdict of the batch lands, before the
next is chosen, so results are consumed in submission order and
every round sees the same state. concurrency: 1 schedules the
same way without the flag, so two same-seed serial runs in one
process return the same rows. Same seed, same concurrency, same
agent gives the same rows; a slow rollout holds its batch, so
uneven latency costs throughput. It needs the clock off, since a
clock stop lands wherever the run happens to be.
Without the flag, which rows land before the cap depends on thread
timing.