Skip to main content
Every name on withwhile.com/platform/runs comes from your code, and the page marks it so. The coding agent that writes the track(...) calls decides what the team reads for months. One rule covers it: read the repo, then use its words. The package name, the prompt files, the policy doc, the existing test names, the deploy tags and the changelog already hold the vocabulary. The SDK’s example names (refund-agent, refund_policy, v1) are for the playbooks, not for a real repo. Three checks before the first post:
  • A teammate who has never opened the platform would recognise the agent id.
  • Every behavior name appears, in some form, in the repo’s own docs or tests.
  • Versions sort the way the team’s releases sort.
Why the test is the exception: a score is comparable only with the setup held constant (rlhfbook.com, “Evaluation”), so the test’s name is its content and changes on its own when an ask changes. Everything else is a label a person reads, and the person is on the team. The AGENTS.md block that whileai init writes carries this in one line: you know this repo best; name the agent, behaviors, versions and experiments in its words, and the test by its content.
Last modified on September 20, 2026