The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/community/who-protects-the-holdout before running the commands below. Browse this recipe on GitHub.decontaminate() applies four rules in order. The first, same_task, compares
scenario_id / task_id rather than words. On a holdout that simulate() wrote,
the call drops 98 to 100% of the training rows (0.981 to 1.000, three seed pairs of
about 106 rows), and same_task is 103 or 104 of every 105 it drops. Strip the ids,
which is what every evaluation set you did not write with simulate() looks like, and
the same call on the same rows drops 18 to 30% (0.178 [0.112, 0.252] to 0.295
[0.219, 0.381]). The report does not say which of those two situations you are in.
What you will learn: which rule is actually carrying your decontamination pass, why an
external eval set is a different regime from an SDK-native one, and what
contamination_rate: 0.0 does and does not certify. You need nothing: no key, no
model, no GPU, no network. Under a minute.
The question
A previous community run (how-much-contamination-survives,
#482) measured the text rules against human-labelled paraphrases (QQP, PAWS) and
found the default lexical rule removes 0.087 [0.076, 0.097] of real contamination. That
is a number about words. It left an obvious hole, which that run named in its own “next”
section:
whetherIt is most of it. And that turns out to be the more useful way round, because it means the 9% number applies exactly where the ids are absent, and nowhere else.same_tasksilently rescues SDK-generated rows, i.e. how much of a realsimulate()holdout is protected by ids rather than by text, because if it is most of it, the text rules matter less than this run implies.
Run it
The setup
Twosimulate() runs from the same seeded agent and the same brief, at different seeds:
one is “training”, one is “holdout”. This is the case the decontaminate docstring names
as the one word overlap cannot see: “a holdout written by re-running the generator on
the same briefs”. Then the same call twice:
- native:
decontaminate(train, against=holdout), rows exactly assimulate()returned them. - id-less: the identical rows, with
scenario_idandtask_idremoved from the evaluation set only.
scenario_ids in seed 0 recur in seed 1), so 104 of
the 107 training rows carry an id the holdout also carries. In the native regime the
rate is that id match, almost by definition. The informative number is the id-less one:
how many of those same rows the text rules find once the ids are gone.
Results
whileai 0.86 (PyPI, fresh venv) · offline, simulator=False, seeded agent · no key,
no GPU, $0. Re-run on whileai 0.88 for review: every count identical, 22 s.
Intervals are wai.pass_at(...).ci95, a percentile bootstrap with one task per row,
which is the bootstrap of a proportion; the exact binomial quantiles on the same counts
agree to 0.01. The 105/105 cell is degenerate under a bootstrap (every resample is 1.0),
so its Wilson interval is given instead.
The native and id-less intervals do not come close to overlapping in any trial. In every
native trial
same_task alone accounts for 103 or 104 of the 105 rows dropped; the text
rules contribute 1 or 2 (all near hits). In the id-less regime the drops are 14 to 17
exact and 5 to 14 near. The scenario_id is robust to the seed because it keys the
situation, not the sampling. That is the SDK working exactly as its docstring says, and
it is good design.
The part to be careful about
Both reports look the same. Here is the id-less one against a real GSM8K holdout (frombuild_sets.py below):
notes is empty. n_same_task: 0 is indistinguishable from “the ids were
compared and none matched”. Filed as #488; it is the complement of #480. Since
that fix the report carries rules_skipped ({"same_task": "0 of 200 evaluation rows carried a scenario_id or task_id"} here) and a notes line saying only the text rules
ran, so the two regimes no longer print the same shape.
The practical rule: if your eval set came from simulate(), the default is strong
and same_task is why. If it came from anywhere else (GSM8K, a Hub set, logged
production traces) you are running on the text rules alone, and their measured recall
is 0.18 to 0.30 here and 0.087 against human-labelled paraphrases (#482). Pass an
embedder= in that case.
The GPU half: what the survivors do to a number
inflation_modal.py and build_sets.py carry the second question: not exposure (how
much leak survives) but inflation (what the survivors do to a measured held-out score).
GSM8K, Qwen/Qwen2.5-1.5B-Instruct, LoRA SFT on one A10G, two arms that differ only in
which decontamination rule cleaned the training set. This half needs datasets,
sentence-transformers, the GSM8K download from the Hub, and your own Modal account.
build_sets.py plants 60 rule-paraphrased copies of held-out questions, with their gold
solutions, into an 800-row GSM8K training pool, then cleans it twice. One deterministic
run at SEED=0; the intervals are exact binomial on the counts:
Base noise floor: the untrained model on the same 200 held-out questions at temperature
0.7, three evals per run. It was measured twice, on two separate A10G containers, which
makes it a replication rather than a single triple:
Six evals, all in [0.490, 0.555]; pooled 632/1200 = 0.527 [0.498, 0.555] Wilson,
sample sd 0.023 across evals, band 0.065. A single 200-question eval carries an interval
of about ±0.07 on its own, so the band is the sampling noise of one eval, not drift. The
two runs’ means differ by 0.007, inside the within-run spread, so the floor is stable
across containers. Any arm-to-arm difference has to clear that band to mean anything.
This half did not finish inside the session. The arm pass rates are not in this
directory. See What did not work. The scripts run as written and
the setup above is reproducible today; the two pass rates are the missing cell.
Note on the 0.333: rule-based paraphrase is a generous input for the lexical rule (it
keeps more n-grams than a human rewording; #482 measured 0.087 on human-labelled
pairs). So 40 surviving leaks is a conservative contamination load, and any inflation
measured from it is a lower bound.
What did not work
- I killed my own GPU run, twice, the same way. Two
modal runinvocations of the same script were live at once (the first launch succeeded although its shell reported an error, so I did not know it existed). They share an app name, somodal app stopon the one I believed was the orphan terminated the other one’s runner too, eight steps into the first arm. I relaunched; that container’s stdout had not reached my log by the time the session’s clock ran out, so I stopped it for the spend guardrail, and its base eval turned out to have completed and flushed a moment earlier. Hence two noise floors and no arms. If you run this:modal app listbefore you stop anything, and give each run a distinctappname. wai.pass_at(rows, k=1).ci95returnsNonewhen every row shares a task id, with no warning and a valid-lookingpass_at_1beside it. It groups by task, so a proportion over rows needs one group per row;proportion()inrun.pydoes that. There is no public binomial-interval helper;pass_atis the only door to one. Filed as #490.- Paraphrasing without a model key. No
ANTHROPIC_API_KEYorOPENAI_API_KEYthis session, so the leaks inbuild_sets.pyare rule-based rewrites (name swaps, phrase substitutions, question moved to the front), not model paraphrases. Their lexical distance is a design choice of mine, which is why the 0.333 above is reported as a property of these paraphrases rather than of paraphrase in general.
Reruns
Everything in the regimes table:results_regimes.json in this directory is the output of that exact command.
results.json is the hand-assembled summary, including the GPU half.
What this does NOT show
No claim about inflation is made here. Exposure is not inflation, and without the arms there is only the former. The id-less numbers (0.18 to 0.30) are specific to this template-driven generator, which hands the text rules a harder input than free-form prose would; the human-labelled comparison is 0.087 (#482).Next
The arm pass rates.modal run inflation_modal.py is about 25 minutes on one A10G and
turns the exposure number into an inflation number: the delta between the two arms on the
same 200 held-out questions, paired, against the 0.065 base noise band. The mechanism
check is whether the gap concentrates on the 40 questions whose paraphrase survived into
the default arm and not on the other 160; leaked_flags in the results file is there
for that split.