FG—01AI Design Field Guide

Search the field guide

Search chapters, principles, patterns, worksheets, and glossary terms

PR—04

Keep the human senior

The AI drafts, proposes, and executes — the human directs, approves, and can always grab the wheel.

The best mental model for agentic AI is a talented, fast, occasionally overconfident junior colleague. You’d never tell that colleague “push whatever you think is right to production,” and no interface should either. Keeping the human senior isn’t a philosophical stance about human dignity — it’s the practical observation that the human holds the context, the accountability, and the consequences, so the human must hold the decision rights.

Seniority is enforced structurally, not by politeness. Claude Code asks before running commands or editing files, and lets you widen its permissions deliberately — per-command, per-session — so autonomy is something you grant, not something the agent assumes. Cursor runs terminal commands through the same gate. Gmail and Notion AI produce drafts that sit inert until you accept them; the send button stays yours. In every case the pattern is the same: the AI proposes, a human disposes, and the cost of the AI being wrong is capped at the cost of reviewing a proposal.

The gate has to be honest work, not a ritual. An approval dialog that shows “Run 3 commands?” without showing the commands trains people to click yes — and an approval that people click reflexively is worse than no approval, because it launders the AI’s decision through a human signature. Show the diff, the recipients, the dollar amount. If reviewing the proposal takes longer than doing the task, your gate is at the wrong altitude.

Seniority also means interruptibility. A senior partner can say “stop, wrong direction” at any moment, and products like ChatGPT and Claude make stopping generation a single click, with agent tools increasingly letting you inject steering mid-run rather than forcing kill-and-restart. An AI you can only cancel is a vending machine; an AI you can redirect is a collaborator. The difference is the entire feel of the product.

And when the system is out of its depth, escalation should go up the hierarchy, not sideways into confabulation. Customer-support AI that hands off to a human with full conversation context intact respects the org chart; AI that improvises a refund policy does not. The rule generalizes: the moment stakes exceed the autonomy you were granted, come back to the senior partner. That’s not a limitation of agentic products — it’s what makes granting them autonomy rational in the first place.

Steelman the other side, because it’s winning funding rounds: the entire economic promise of agents is that the human is not in the loop. An assistant that interrupts you for every file edit has merely converted your work into supervision — often slower than doing the task yourself — and approval fatigue is a real, measured phenomenon: gates people click through reflexively provide the liability of oversight without the substance. Meanwhile competitors ship “fully autonomous” agents that feel like magic in demos, and the market rewards the feeling. The rebuttal is that the argument confuses the destination with the route. Autonomy is earned per task type, not declared per product — Claude Code’s permission model exists precisely so trusted operations graduate out of the gate while consequential ones stay in it. The products that removed the human wholesale are the ones generating the incident postmortems; the ones that made seniority granular are the ones people hand bigger jobs every month. Supervision that shrinks is the feature.

The measurement is the health of the approval gate itself. Track approval dwell time on consequential actions: if median time-to-approve collapses toward reflexive (under a second or two) while the actions being approved remain diverse and high-stakes, your gate has become a rubber stamp and is actively laundering risk. Track the rejection-and-modification rate at gates — a gate nobody ever rejects at is either guarding a trusted operation that should graduate to autonomy, or it’s theater; a healthy gate shows a persistent single-digit rejection rate, proof it sits where judgment still matters. For interruptibility, measure mid-run steering: what share of corrections happen via interrupt-and-redirect versus cancel-and-restart, and how often escalations to a human arrive with context intact versus forcing the user to re-explain. Good looks like autonomy expanding measurably over a user’s tenure — more actions auto-approved by explicit grant — while incidents traced to unapproved consequential actions hold at zero. That pairing is seniority working.