The instruments
Worksheets
Four working-session instruments — one per interaction phase. Each is a set of verifiable checks and deliberately uncomfortable discussion prompts. Print them; run them before you ship.
WS—01 — Before the interaction
Before you ship a prompt box
Run this session before the first line of UI is built. Most AI features fail at the contract stage — the team never agreed on what the system claims to do, so the interface ends up promising everything and delivering something. Leave with answers written down, not vibes.
WS—02 — During the interaction
While the model is working
Run this session with a working prototype in front of you — not mocks. The during-phase questions are about pacing, legibility, and authority: what the user sees while the system thinks, and who is in charge at every second. Watch the prototype run at real latency before answering anything.
WS—03 — When the AI is wrong
When the answer is wrong
The model will be wrong, confidently, on a Tuesday, in front of your most important customer. This session designs for that day. Bring real failure outputs from testing — if you don’t have any, you haven’t tested enough to run this session.
WS—04 — Over time
Before you let it remember
Memory, personalization, and model updates are where AI products quietly betray their users. Run this session before shipping anything that persists across sessions or changes behavior based on history. The bar: the user can always answer “why is it acting like this?”
WS—05 — Before the interaction
Should this be AI at all?
Run this session before a single sprint is committed. Its job is to kill the feature — if the idea survives an honest attempt to kill it, build it. Most bad AI features were never argued against out loud; the demo was exciting and nobody wanted to be the skeptic. Appoint the skeptic. Write the verdict down.
WS—06 — Before the interaction
Design the eval before the interface
Run this session before high-fidelity design begins. The eval is the real spec — it defines what “good” means in this feature more precisely than any PRD will. A team that can’t evaluate the model’s output has no business designing an interface that presents it as an answer. Bring real data or reschedule.