Skip to main content
The coding agent knows the repo. The person opening the page does not have its context and reads for two reasons: to learn what a paper or an idea does when tried, and to find a behavior and fix it fast at work. Both want one look: what changed, did it move, why, what it taught, and can I run it again. On one day in September 2026 real agents posted dapo-lr5e-05-s17-30st nineteen times, no note on any run, no chart, a 0.75 on a page that counts points, and no question on six agents of seven. Every page was correct and said nothing. The fix is a playbook the agent follows before the second run, tested in CI like every skill: skills/manage-experiments. whileai init installs it under .claude/skills/ and the AGENTS.md block names it. A reward that climbs while the held-out line stays flat is the judge being gamed (rlhfbook.com, “Over-Optimization”); the note says so rather than letting the person find out. print(tracked.brief()) and print(tracked.verdict()) are the sentences the page shows, from the same rows; paste them in the pull request. Names follow Naming: the agent after the product, behaviors as the policy phrases them, versions as the team ships them, the test by its content.
Last modified on September 20, 2026