dapo-lr5e-05-s17-30st nineteen times, no note on any run, no chart, a 0.75 on a page that counts points, and no question on six agents of seven. Every page was correct and said nothing.
The fix is a playbook the agent follows before the second run, tested in CI like every skill: skills/manage-experiments. whileai init installs it under .claude/skills/ and the AGENTS.md block names it.
A reward that climbs while the held-out line stays flat is the judge being gamed (rlhfbook.com, “Over-Optimization”); the note says so rather than letting the person find out.
print(tracked.brief()) and print(tracked.verdict()) are the sentences the page shows, from the same rows; paste them in the pull request. Names follow Naming: the agent after the product, behaviors as the policy phrases them, versions as the team ships them, the test by its content.