WS—06 · Before the interaction
Design the eval before the interface
Run this session before high-fidelity design begins. The eval is the real spec — it defines what “good” means in this feature more precisely than any PRD will. A team that can’t evaluate the model’s output has no business designing an interface that presents it as an answer. Bring real data or reschedule.
Checks
Discussion prompts
- If the eval and a stakeholder’s gut disagree about whether an output is good, which one wins — and has that ever actually happened here?
- Who wrote our grading criteria, and have two people independently graded the same ten outputs to see if they agree?
- What behavior are we not measuring because it is hard to measure, and what happens if that’s the thing users notice first?
- When the ten worst outputs went up on the screen, who in the room was surprised — and why were the rest of us not sharing what we knew?
- Would this feature pass its own eval today? If nobody can answer in one sentence, what is the interface being designed against?
- What is the smallest prompt change that ever shipped without a rerun of the golden set, and how would we even know?