The mechanism
The score is called the reward. There are two ways to get one.- A program. The answer matches the reference, the tests pass, the tool was called before the reply. When a program can check the job, use the program. It is cheap, repeatable, and cannot be flattered. The field calls this a verifiable reward, and the program a verifier.
- A model. When no program can check the job (was the tone right, did it explain the policy), a model reads the reply against a written rubric and answers 0 or 1. That is an LLM judge.
Run it
The setup is lesson 2’s run. Then three rewards on the same rows.The numbers after each pass rate are the interval. Lesson 4 is about
them. For now:
0.67 [0.55..0.78] means the true rate very likely sits
between 55% and 78%.Where it comes from
- Lambert, N. et al. Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv:2411.15124, 2024. Reinforcement learning with verifiable rewards: the reward is a program.
- Zheng, L. et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS, 2023. arXiv:2306.05685. How well a model judge agrees with people, and where it is biased.
- Cohen, J. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1), 1960. Kappa.
- Lambert, N. Reinforcement Learning from Human Feedback. arXiv:2504.12501, 2025. Chapter Reward Modeling.
Next
A pass rate without an interval is a guess: what0.67 [0.55..0.78] means and why the brackets matter more than the
number.