I just posted this on LinkedIn to argue that AI safety conversations over-index on preventing mistakes and underinvest in making them cheap to undo. The core idea: for reversible tasks, optimize for trivial reversion rather than strict pre-approval.
I don't have feedback yet, but the immediate question that typically follows is: Okay, but how? A 30% undo rate is uninterpretable on its own. It could mean your model is struggling, your users are exploring, or your instructions are unclear. Here's how you'd tease those apart.
The Core Problem: Raw Undo Rate Is Noise
Tracking reversions is easy. Interpreting them is hard. A user hitting "undo" compresses multiple failure modes into a single binary event. Without structure, you're flying blind.
1. Distinguish Undo Types
Not all reversions indicate model failure. You need to tag the reason:
- Correction: Output was wrong or low-quality (the signal you want)
- Preference: Output was fine, user just wanted something different
- Exploration: User was trying options, not evaluating quality
- Cascading: Undoing because something downstream broke
Only Correction cleanly indicates model failure. The others are noise for quality measurement (though still useful for understanding usage patterns).
How to capture this:
- Lightweight: One-click reason selector on undo (takes 2 seconds)
- Passive: Infer from behavioral signals. A reversion after 3 seconds and zero edits? Likely Correction. After 5 minutes and multiple tweaks? Probably Preference.
2. Measure Time-to-Undo
Speed matters. The distribution tells you different things:
- <10-30 seconds: Obvious errors. User glanced and immediately rejected.
- 2-3 days: Delayed consequences. Subtle problems that surfaced in use, or changed requirements.
- Flat distribution: Likely Preference/Exploration noise.
A spike at under 10 seconds means your model is producing visibly bad outputs. A spike at 2-3 days means you're hitting the caveat I mentioned—undo doesn't help with downstream failures.
3. Track Undo Depth, Not Just Occurrence
What happens after the undo matters more than the undo itself:
- Undo → manual rewrite: Model couldn't do the task
- Undo → re-prompt → accept: Model could do it, instructions were unclear
- Undo → re-prompt → undo again → give up: Task likely outside model capability
The recovery path pinpoints where the failure lives: model, prompt, or task design.
4. Build a Comparison Baseline
Without a baseline, "30% undo rate" is meaningless. Is that good or terrible? Compare:
- Across task types: Where does the model struggle most?
- Across model versions: A/B test new versions against undo rate
- Against human baseline: What was the revision rate before AI? If human drafts had a 25% "undo" equivalent, 30% might be acceptable.
Context transforms a number into a decision.
5. Sample for Ground Truth (Catching Automation Bias)
Undo tracking won't catch cases where users should have undone but didn't. This is automation bias—blindly trusting the AI.
Periodically audit accepted outputs:
- Randomly sample 5% of "accepted" work
- Have a human expert evaluate quality
- Compare audit failure rate to undo rate
If audits find 20% of accepted outputs are poor but only 5% get undone, you've quantified your automation bias gap. That's your ceiling on how much autonomy you can safely grant.
6. Instrument the Non-Undos
Tracking what didn't get undone is as important as tracking what did:
- Review time: How long before accepting? If average review time drops 15% month-over-month, that's early warning of automation bias.
- Edit rate: Users who accept then make small tweaks are still engaged. Users who accept with zero changes might not be reading at all.
These metrics catch complacency before it causes damage.
What You'd Actually See: A Useful Dashboard
A naive dashboard shows:
Undo Rate: 30%
That's useless. A useful dashboard shows:
Total Reversions: 30%
├─ Correction (quality signal): 18%
├─ Preference/Exploration (usage pattern): 12%
Time-to-Undo Distribution:
├─ <30s (obvious errors): 55%
├─ 30s-24h (prompt refinement): 30%
└─ >24h (delayed consequences): 15%
Recovery Paths:
├─ Undo → Manual Rewrite: 8%
├─ Undo → Re-prompt → Accept: 19%
└─ Undo → Re-prompt → Abandon: 3%
Automation Bias Audit:
├─ Sampled Accept Rate Flagged as Poor: 8%
└─ Estimated True Error Rate: 26%
Review Time Trend: ↓ 15% over past month (warning)
This tells you where the model struggles, why users revert, and how much bias is creeping in.
The Honest Caveats
This isn't free. The automation bias audit requires human labor—expensive and hard to maintain at scale. Passive inference of undo reasons works until it doesn't, and you'll need to validate those signals. Most importantly, this framework only applies where reversibility is feasible. For patient-facing decisions or irreversible actions, none of this replaces traditional safety measures.
But for internal workflows where you can walk things back, instrumenting reversibility gives you richer, more honest signal than most pre-deployment evals. It's not a paradigm shift. It's just an underweighted factor that deserves more weight than it gets.
I write about putting AI in front of people who do real work. I'm moving to Paris as a Forward Deployed Engineer. Read about my work, read more posts, or write to me.