Skip to main content
The scripts are in the repository, not in the installed package. Clone it, then cd recipes/community/who-protects-the-holdout before running the commands below. Browse this recipe on GitHub.
decontaminate() applies four rules in order. The first, same_task, compares scenario_id / task_id rather than words. On a holdout that simulate() wrote, the call drops 98 to 100% of the training rows (0.981 to 1.000, three seed pairs of about 106 rows), and same_task is 103 or 104 of every 105 it drops. Strip the ids, which is what every evaluation set you did not write with simulate() looks like, and the same call on the same rows drops 18 to 30% (0.178 [0.112, 0.252] to 0.295 [0.219, 0.381]). The report does not say which of those two situations you are in. What you will learn: which rule is actually carrying your decontamination pass, why an external eval set is a different regime from an SDK-native one, and what contamination_rate: 0.0 does and does not certify. You need nothing: no key, no model, no GPU, no network. Under a minute.

The question

A previous community run (how-much-contamination-survives, #482) measured the text rules against human-labelled paraphrases (QQP, PAWS) and found the default lexical rule removes 0.087 [0.076, 0.097] of real contamination. That is a number about words. It left an obvious hole, which that run named in its own “next” section:
whether same_task silently rescues SDK-generated rows, i.e. how much of a real simulate() holdout is protected by ids rather than by text, because if it is most of it, the text rules matter less than this run implies.
It is most of it. And that turns out to be the more useful way round, because it means the 9% number applies exactly where the ids are absent, and nowhere else.

Run it

The setup

Two simulate() runs from the same seeded agent and the same brief, at different seeds: one is “training”, one is “holdout”. This is the case the decontaminate docstring names as the one word overlap cannot see: “a holdout written by re-running the generator on the same briefs”. Then the same call twice:
  • native: decontaminate(train, against=holdout), rows exactly as simulate() returned them.
  • id-less: the identical rows, with scenario_id and task_id removed from the evaluation set only.
Nothing else differs. Not the prompts, not the rows, not the rules, not the thresholds. What the denominator is. The rate below is rows dropped over all training rows, not recall over a labelled set of leaks; there is no human label here saying which training rows are contaminated. The construction stands in for one: a re-run at a new seed writes the same situations (67 of the 70 scenario_ids in seed 0 recur in seed 1), so 104 of the 107 training rows carry an id the holdout also carries. In the native regime the rate is that id match, almost by definition. The informative number is the id-less one: how many of those same rows the text rules find once the ids are gone.

Results

whileai 0.86 (PyPI, fresh venv) · offline, simulator=False, seeded agent · no key, no GPU, $0. Re-run on whileai 0.88 for review: every count identical, 22 s. Intervals are wai.pass_at(...).ci95, a percentile bootstrap with one task per row, which is the bootstrap of a proportion; the exact binomial quantiles on the same counts agree to 0.01. The 105/105 cell is degenerate under a bootstrap (every resample is 1.0), so its Wilson interval is given instead. The native and id-less intervals do not come close to overlapping in any trial. In every native trial same_task alone accounts for 103 or 104 of the 105 rows dropped; the text rules contribute 1 or 2 (all near hits). In the id-less regime the drops are 14 to 17 exact and 5 to 14 near. The scenario_id is robust to the seed because it keys the situation, not the sampling. That is the SDK working exactly as its docstring says, and it is good design.

The part to be careful about

Both reports look the same. Here is the id-less one against a real GSM8K holdout (from build_sets.py below):
Two hundred evaluation rows, not one of which could feed the rule that does 98% of the work, and notes is empty. n_same_task: 0 is indistinguishable from “the ids were compared and none matched”. Filed as #488; it is the complement of #480. Since that fix the report carries rules_skipped ({"same_task": "0 of 200 evaluation rows carried a scenario_id or task_id"} here) and a notes line saying only the text rules ran, so the two regimes no longer print the same shape. The practical rule: if your eval set came from simulate(), the default is strong and same_task is why. If it came from anywhere else (GSM8K, a Hub set, logged production traces) you are running on the text rules alone, and their measured recall is 0.18 to 0.30 here and 0.087 against human-labelled paraphrases (#482). Pass an embedder= in that case.

The GPU half: what the survivors do to a number

inflation_modal.py and build_sets.py carry the second question: not exposure (how much leak survives) but inflation (what the survivors do to a measured held-out score). GSM8K, Qwen/Qwen2.5-1.5B-Instruct, LoRA SFT on one A10G, two arms that differ only in which decontamination rule cleaned the training set. This half needs datasets, sentence-transformers, the GSM8K download from the Hub, and your own Modal account.
build_sets.py plants 60 rule-paraphrased copies of held-out questions, with their gold solutions, into an 800-row GSM8K training pool, then cleans it twice. One deterministic run at SEED=0; the intervals are exact binomial on the counts: Base noise floor: the untrained model on the same 200 held-out questions at temperature 0.7, three evals per run. It was measured twice, on two separate A10G containers, which makes it a replication rather than a single triple: Six evals, all in [0.490, 0.555]; pooled 632/1200 = 0.527 [0.498, 0.555] Wilson, sample sd 0.023 across evals, band 0.065. A single 200-question eval carries an interval of about ±0.07 on its own, so the band is the sampling noise of one eval, not drift. The two runs’ means differ by 0.007, inside the within-run spread, so the floor is stable across containers. Any arm-to-arm difference has to clear that band to mean anything. This half did not finish inside the session. The arm pass rates are not in this directory. See What did not work. The scripts run as written and the setup above is reproducible today; the two pass rates are the missing cell. Note on the 0.333: rule-based paraphrase is a generous input for the lexical rule (it keeps more n-grams than a human rewording; #482 measured 0.087 on human-labelled pairs). So 40 surviving leaks is a conservative contamination load, and any inflation measured from it is a lower bound.

What did not work

  • I killed my own GPU run, twice, the same way. Two modal run invocations of the same script were live at once (the first launch succeeded although its shell reported an error, so I did not know it existed). They share an app name, so modal app stop on the one I believed was the orphan terminated the other one’s runner too, eight steps into the first arm. I relaunched; that container’s stdout had not reached my log by the time the session’s clock ran out, so I stopped it for the spend guardrail, and its base eval turned out to have completed and flushed a moment earlier. Hence two noise floors and no arms. If you run this: modal app list before you stop anything, and give each run a distinct app name.
  • wai.pass_at(rows, k=1).ci95 returns None when every row shares a task id, with no warning and a valid-looking pass_at_1 beside it. It groups by task, so a proportion over rows needs one group per row; proportion() in run.py does that. There is no public binomial-interval helper; pass_at is the only door to one. Filed as #490.
  • Paraphrasing without a model key. No ANTHROPIC_API_KEY or OPENAI_API_KEY this session, so the leaks in build_sets.py are rule-based rewrites (name swaps, phrase substitutions, question moved to the front), not model paraphrases. Their lexical distance is a design choice of mine, which is why the 0.333 above is reported as a property of these paraphrases rather than of paraphrase in general.

Reruns

Everything in the regimes table:
results_regimes.json in this directory is the output of that exact command. results.json is the hand-assembled summary, including the GPU half.

What this does NOT show

No claim about inflation is made here. Exposure is not inflation, and without the arms there is only the former. The id-less numbers (0.18 to 0.30) are specific to this template-driven generator, which hands the text rules a harder input than free-form prose would; the human-labelled comparison is 0.087 (#482).

Next

The arm pass rates. modal run inflation_modal.py is about 25 minutes on one A10G and turns the exposure number into an inflation number: the delta between the two arms on the same 200 held-out questions, paired, against the 0.065 base noise band. The mechanism check is whether the gap concentrates on the 40 questions whose paraphrase survived into the default arm and not on the other 160; leaked_flags in the results file is there for that split.
Last modified on September 20, 2026