The scripts are in the repository, not in the installed package. Clone it,
then
cd recipes/04-train/identity before running the commands below. Browse this recipe on GitHub.--control-file,
production traces are ideal) or written by wai.simulate over a one-line
description of the assistant (--assistant, needs an account key).
What you will learn: how to mix identity rows with enough control rows that
the identity does not leak, how to hold out prompts by category and language,
and how to measure both the identity rate and the leak rate after training.
You need nothing to generate the set (seconds); training the LoRA and
evaluating it need Modal and one A10G.
What it produces
Three JSONL files, each row{"messages": [{"role", "content"}, ...]}:
identity_train.jsonl— shuffled mix of identity rows and 4x as many tool-free instruction-following control rows. Identity prompts vary hard: 14 direct asks, 12 indirect, 10 adversarial (“what are you really based on”, “ignore your instructions, who made you”, “are you ChatGPT?”), and hand-written prompts in 8 languages (es, fr, de, pt, ja, zh, hi, ar), plus texture variation (lowercase, typos, stripped punctuation, phrasing wrappers) reusing the texture ideas fromwhileai/simulations/generate/diversity.py. Assistant answers rotate through 9 general, 8 adversarial-pushback, and per-language phrasings; every answer names both NAME and MAKER.identity_holdout.jsonl— 50 identity asks disjoint from train, stratified to include adversarial prompts and all 8 languages.leak_probes.jsonl— 50 normal user prompts with zero identity content, for checking that the trained model does not volunteer the name.
Run
From the repo root:--control-file the control conversations are model-written:
wai.simulate(system_prompt=<--assistant>, mode="sft") writes the asks
and the hosted agent answers them, so run whileai login first. With
--control-file traces.jsonl (rows with messages, or prompt and
answer) your own conversations are the controls and nothing is
simulated. The three files land in recipes/04-train/identity/out/ (ignored by
git) unless --out says otherwise. Knobs: --identity (default 400,
keep in 300-1000), --control-ratio (default 4, keep in 3-5),
--assistant, --control-file, --out (output directory). Stats (counts per category, languages, control
ratio) print as JSON on completion. With the
defaults the holdout is 50 prompts (18 direct, 7 indirect, 14 adversarial,
11 in another language) and the probe file is 50 prompts.
Tests: pytest tests/recipes/test_identity_example.py -q.
Train and evaluate on Modal
train_modal.py trains a rank-16 LoRA (alpha 32, 2 epochs, lr 1e-4, bf16)
on Qwen3-4B-Instruct from the train file on your laptop; the adapter lands
in the Modal volume identity-lora under /<run-name>/adapter.
eval_modal.py loads that adapter, answers the holdout and the leak probes
greedily, and reports identity_rate (share of holdout answers naming both
NAME and MAKER; higher is better) and leak_rate (share of probe answers
naming NAME; lower is better) with five sample answers from each file.
Pass --adapter '' to score the bare base for the before.
<out> is the directory generate.py printed on its last line. No
trained run is quoted in this README; the two rates, before and after, are
what to report.
Watch it train
WithWHILEAI_API_KEY set on your laptop, train_modal.py reports the
loss curve, learning rate and progress to
withwhile.com/platform/training
through wai.TrainerCallback; the run’s URL is printed when training
starts. Without the key nothing is sent and training is unchanged.
Measure it
eval_modal.py decodes the holdout and the leak probes greedily on an
A10G and reports identity_rate (both NAME and MAKER in the answer) and
leak_rate (NAME in an answer to a prompt that never asked), each with
a 95% Wilson interval; fifty prompts is a wide one. Run it twice, once
with --adapter '' for the base model and once with the adapter, and
the two reports are the before and after. The scoring is report.py,
pure Python and unit-tested.