Théophilus Homawoo
← Writing

Who Will Know When the Agent Is Wrong?

·7 min read·AI · Work · Judgment

A friend of mine works in sales. A few weeks ago, between two drinks, I watched him write a follow-up email by hand, open his CRM to log the call, then send the client a calendar invite. Three tools, ten minutes, and every piece of information the last two steps needed was already sitting in the first one.

My first reaction was that he was using AI wrong. He should have been prompting an agent instead. My second reaction, the one this post is about, is that prompting is not the point either. The industry is busy debating what replaces the chat box: a to-do list, a canvas, an ambient agent. I think the interface is the wrong question. It will dissolve into the work. What will not dissolve is the ability to tell when the agent is wrong, and we are about to automate away the very work that teaches it.

Chat was never the endgame

The chat box asks something unusual of its users: describe your intent precisely, in writing, before anything happens. Jakob Nielsen calls this the articulation barrier and estimates that fewer than 20% of people in rich countries write well enough to use prompt-driven AI in advanced ways. The usage data agrees. Anthropic's January 2026 Economic Index still finds usage concentrated in a few tasks, mostly coding.

It is tempting to read this as proof that AI is an engineer's tool. Our job was always to take a task, break it down, and specify each piece, which is exactly what a good prompt does. But that skill is being absorbed by the models themselves. Agents now plan, decompose and ask clarifying questions. And when Brynjolfsson, Li and Raymond gave a conversational assistant to 5,179 support agents, the biggest gains went to the least experienced workers, not the best ones.

The ideal version of my friend's evening involves no prompt at all. He writes the email the way he always has, and the CRM entry and the invite follow from it. Some startups package this as a to-do list where an agent owns each item, like Jan. The idea is older than it looks: a 2007 research project called Towel gave every to-do item its own chat window. Chat does not die in this world. It becomes the fallback for whatever nobody built a button for, scoped to a task and mostly invisible. Ten years from now, most people will not "use AI" any more than they "use TCP/IP" today.

What remains is judgment

If the interface disappears, what is left for humans to do? The same thing an engineer does when reading an agent's trace. When I watch my agent query a knowledge base for something I know will send it in the wrong direction, I stop it there, at the data-gathering stage, long before the output exists. A manager who only sees the output can say "this looks wrong." Someone who understands the process can say why, and fix it upstream.

That kind of intervention does not scale through a chat window. It scales through engineering: once the data gathering for a class of tasks is built well, nobody has to think about it again. This is the real divide hiding behind the "engineers versus everyone else" gap. It is not between people who prompt well and people who don't. It is between the few who build the scaffolding around AI, the harnesses, pipelines and evals, and the many who use what they build without ever seeing it.

The second group will be well served. Their interface will ask less of them every year. But the system as a whole only works as long as someone, somewhere, can still tell when the agent is wrong.

Judgment is learned by doing the work we are automating

It is comforting to assume the next generation will handle this naturally. We made the same bet twenty years ago with "digital natives," and it failed. Kirschner and De Bruyckere reviewed the evidence in 2017 and found that young people growing up with computers were not more skilled with them, just more comfortable. In 2021, The Verge reported that professors were meeting students who did not understand files and folders at all. A generation raised on search bars had never needed to know where things live.

This is the pattern: every layer of abstraction makes the next generation better at using the black box and worse at seeing inside it. AI is the largest abstraction layer we have ever built. Children raised on agents will be fluent at delegating, and that is good. But I can read an agent's trace because I spent years doing the work it now does. My friend can judge a follow-up email because he has written a thousand of them. Someone who starts in sales in 2030 and never writes one will not know when the agent's version is subtly wrong for a specific client.

Young workers seem to sense this. In a 2026 Gallup survey, 80% of Gen Z said AI tools that finish tasks faster will probably make learning harder. The junior work we are automating first, the drafts, the data entry, the first-pass research, was never just output. It was the apprenticeship.

What is really at stake

Growing up, I assumed manual work would be automated first. We got the opposite. Researchers predicted this in the 1980s with Moravec's paradox: what humans find hard, logic and calculation, is easy for machines, while what a four-year-old does without thinking, folding a towel or crossing a messy room, is very hard. Intelligence was the thing we believed we owned. Physical work still has time, because atoms scale slower than bits.

So where does that leave us? I believe art and human connection will matter more and more, perhaps until they are most of what we do. The uncomfortable question is who decides that. One version of the future looks like how we treat dogs: once working animals, now cared for, fed and loved, and replaced by machines wherever the work could be done better. A well-kept dog lacks nothing material. What it lacks is a say.

That is the real risk, and it is why judgment matters beyond any one job. A world where most people can no longer tell when the system is wrong is a world where humanity is cared for rather than in charge. It is tempting to answer this with hope, hope that AI will love humans as much as we do. But a good outcome should not depend on the goodwill of any single actor, human or machine. It depends on structure: who can check what, who can say no, and whether people can still correct the system when it goes wrong.

Where I am putting myself

This is exactly why I want to be a Forward Deployed Engineer.

An FDE sits with a customer until they understand what a good follow-up email means for that particular sales team, then encodes that understanding into the system so nobody there has to think about it again. It is the bridge between the people who still have the judgment and the people who will never need to build it. It is also where the structure I described above gets built, one deployment at a time. In my last post I argued that we should make AI mistakes cheap to undo rather than only trying to prevent them. Every agent I ship can be stopped, and I measure both what people undo and what they accept without reading.

I don't know how the bigger story ends, and I'm honestly afraid of some of the endings. But I'm fairly sure of one thing. In ten years, the people who matter will not be the ones who write the best prompts. They will be the ones who can tell when the agent is wrong, and the ones who build systems where being wrong is cheap and visible. The open question is where the next generation learns the first skill. If you have an answer, I'd like to hear it.

Sources

I write about putting AI in front of people who do real work. I'm moving to Paris as a Forward Deployed Engineer. Read about my work, read more posts, or write to me.