PR—09
Feedback must matter
Never ask for feedback you won’t visibly act on — close the loop or drop the button.
Most AI feedback UI is a lie of omission. The thumbs up/down pair suggests a dialogue — you tell the system, the system adjusts — when in reality most of those clicks flow into an offline training pool that will influence a model months from now, if ever. People figure this out within a handful of uses. Watch feedback button engagement decay over a user’s first month and you are watching them learn that the product doesn’t listen. That lesson generalizes; it curdles their read of the whole relationship.
The feedback that works is feedback with a visible landing. When you edit ChatGPT’s memory of your preferences, the very next answer reflects it. When you correct Claude mid-conversation — “stop using bullet points” — the correction takes hold immediately, and the best products persist it rather than making you repeat it every session. Cursor applying your coding conventions after you state them once is feedback closing its loop at the pace people actually expect: now, not next quarter.
This reframes the design task. The question is not “where do we put the rating widget” but “when a person tells us we got it wrong, what changes, and how do they see the change?” Sometimes the answer is an inline acknowledgment — “Noted, I’ll keep summaries under 100 words” — written into legible memory where it can be inspected later. Sometimes it’s an immediate regeneration honoring the correction. What it can never be is a toast that says “thanks for your feedback” followed by identical behavior, which is the interaction-design equivalent of a suggestion box mounted over a shredder.
Implicit feedback deserves the same honesty, applied in reverse. Products quietly learn from which suggestions you accept, which drafts you rewrite, which results you copy — Copilot’s acceptance data being the canonical example. Learning from behavior is fine; laundering it into invisible personalization is not. If implicit signals change the system’s behavior, surface that on the same inspection panel as explicit preferences, or you’ve rebuilt the illegible-memory problem one layer down.
And have the discipline to delete the dead affordances. If a feedback control feeds nothing the user will ever perceive, it is decoration that costs trust with every ignored click. Fewer feedback surfaces that demonstrably matter beat a product wallpapered with inert thumbs. The bar is simple: every way of telling the system about itself should have a visible answer to “what did that just do?”
The strongest case for the humble inert thumb is that slow feedback is how the product actually gets better. Those clicks flowing into an offline pool are RLHF fuel — the aggregate signal that trains the next model — and demanding a visible per-user response from every rating would mean either faking one or building per-user adaptation machinery that most teams can’t ship responsibly. Instant visible adaptation carries its own hazards: a system that reshapes itself on every offhand correction becomes twitchy and manipulable, one sarcastic “great, more bullet points” away from a broken preference, and users often rate outputs poorly for reasons (mood, misread, wrong expectations) that shouldn’t change anything. Stability is a feature; not every signal deserves a landing. Fair — but this argues for honesty about the loop, not for the deceptive widget. Label the slow path for what it is (“help improve the model”), keep the fast path for what the user will feel (“stop doing this”), and give corrections weight proportional to their explicitness. The sin was never collecting training data; it was dressing a donation box as a dialogue.
The decay curve is the measurement. Plot feedback-control engagement across each user’s tenure: in a product where feedback visibly lands, correction usage holds steady or grows with investment; in a suggestion-box-over-a-shredder product, it decays toward zero by week four as people learn the truth. Pair it with correction persistence: after an explicit correction (“keep summaries under 100 words”), sample the next twenty relevant outputs and score compliance — good looks like near-total adherence within the session and across sessions, with the repeat-correction rate (users restating a preference they’ve already stated) driven toward zero, because every repeat is a measured broken promise. For the slow loop, track whether users who give ratings retain better than matched users who don’t; if rating the product predicts nothing about staying, the affordance is spending trust and buying none. And audit your own surfaces quarterly: any feedback control whose signal provably changes nothing a user will ever perceive gets labeled honestly or deleted. The count of those you removed is itself a health metric.