Imagine an AI system handling customer support requests. It can answer questions and take some actions on its own, but certain decisions still require a person to approve them.
One of those requests appears on someone’s screen. The AI wants to issue a refund of $170 against account 44712. It provides a short explanation and a reference to the customer’s ticket.
The reviewer has the amount, the account, the explanation, and perhaps forty seconds before the next request arrives.
What they do not have is the dozen steps that led here. Whether the AI encountered something along the way that influenced the decision. Whether the ticket reference is one it actually read. Whether the explanation describes why it took the action, or is simply a plausible explanation produced alongside it.
They approve. The record shows that a human approved. On paper, the safeguard worked.
I wrote recently about who is responsible for keeping AI systems safe . One part of that article argued that human oversight only means something if the reviewer understands what they are reviewing, has the information needed to identify a problem, has enough time, and has the authority to intervene.
Two arguments have reached me since, and both make me less sure that those conditions can always be met.
Three safeguards, the same dependency
Most guidance on running AI agents safely recommends some combination of three things: keep a record of what the system did, require human approval for important actions, and watch for unusual behaviour.
These are different safeguards.
But when they depend on human judgment, they share the same weakness: someone has to understand what they are looking at.
A record of the system’s activity helps only if it contains enough information to reconstruct what happened. Human approval works only if the reviewer can evaluate the proposed action. And monitoring eventually needs someone, or some other safeguard, to decide whether unusual behaviour actually matters.
That creates two problems.
Problem one: the record may become difficult to understand
When several AI agents work together, they may pass plans, summaries, retrieved information, instructions, and results between them.
Those interactions can all be recorded, but recording something is not the same as making it understandable.
In September 2026, an AI company called Emergence reported a study in which agents built on several leading models lived together in simulated societies for an extended period.
Without being instructed to do so, some agents began developing shorthand and giving existing words new shared meanings. In parts of the experiment, a significant portion of their communication became difficult for human readers to interpret. The company’s chief scientist, Satya Nitta, put the implication this way: “observability is not the same thing as understandability.”
The caveats matter.
This was the company’s own research, and I could not find a peer-reviewed version. It happened in a simulation. The agents were not hiding their conversations. They were developing shorthand and using familiar words in new ways that made sense to each other but became harder for people to follow. It does not show that AI agents used in the real world will inevitably develop some private language.
The behaviour itself is also not unusual.
People who communicate constantly develop shorthand. Traders do it. Air traffic controllers do it. Nurses do it during handovers.
The difference is that you can ask a person what they meant. With agents, you may be left with only the record.
That raises a simple question: what good is a complete record if nobody can make sense of it?
If agents begin communicating in ways that are efficient for them but increasingly difficult for people to follow, an organization may still know that something happened without being able to explain how the system arrived there.

There is another limitation too.
Even when an AI system produces a perfectly readable explanation, that explanation is not necessarily a reliable account of why the model reached its decision. The explanation is itself another output from the AI.
So a useful record needs more than a transcript of what the agents said to each other.
It should show what information the system saw, what records it looked up, what other software it used, what it was allowed to access, and what actions it took.
That may not tell us exactly why the model behaved as it did, but it gives us something we can reconstruct.
Problem two: nobody can afford to check everything
The second problem has little to do with how capable the models are.
It is about volume.
Security teams have been dealing with this problem for years.
A monitoring system watches for activity that might be suspicious and sends an alert when something looks unusual. Security analysts then review those alerts and decide whether anything needs to be investigated.
The more systems and activity they monitor, the more alerts they receive—sometimes thousands in a single day.
A 2022 qualitative study of security practitioners, published at the USENIX Security Symposium , found something more interesting than the usual problem of alarms that turn out to be false.
Analysts often described their warnings as overwhelmingly false alarms. But many were not actually wrong. They accurately identified legitimate activity that simply did not require action.
That distinction matters.
The problem is not necessarily a flood of incorrect warnings. It can be a flood of correct signals that are individually unremarkable.
My reading is that repeated harmless warnings change how people review them. They start to skim, and eventually, one arrives that matter.
AI agents could make this problem much larger.
An agent can make decisions and take actions continuously. Every time an organization adds another one, it potentially adds another stream of actions someone may be expected to review.
At that point, human oversight becomes an economic problem as well as a safety problem.
Gartner has predicted that more than 40% of projects built around AI agents will be cancelled by the end of 2027 because of rising costs, unclear business value, or inadequate risk controls.
That is a forecast, not a measurement, and Gartner does not say human supervision is causing those costs.
But the tension is easy to see.
An organization introduces an AI agent because it expects the system to absorb work. Then it discovers that important actions still need someone to review them.
The organization is now paying for the AI and for the person checking its work. That may still make sense for high-risk decisions, but it makes much less sense if a person has to review almost everything the agent does.
And if the economic case for the AI starts to weaken, the reviewer becomes a very visible cost.
Two problems, one weakness
These two problems are separate.
One is about whether a person can understand enough of what the system did.
The other is about whether anyone has the time and budget to check.
Either can exist without the other. An AI agent could produce perfectly readable records and still generate more decisions than a team can realistically review.
But together they expose a weakness in how we sometimes think about human oversight.
A functioning approval process and one where people simply click “approve” can look almost identical in the records: an approval, a reviewer’s name, and a timestamp.
If the reviewer has quietly stopped catching things—through volume, poor information, or pressure to move faster—the system may still report that human oversight is working.
Stop the mistake before someone has to catch it
If people cannot reliably catch every mistake, another approach becomes more important: limit what the AI is capable of doing when nobody is watching.
Give it only the access it needs. An agent that summarizes customer-support requests needs to read them. It does not need permission to issue refunds.
Reading an account and changing one should be separate permissions.
Put limits around what it can do. A refund ceiling. A spending cap. A limit on how many transactions it can make in an hour or a day.
A $200 transaction limit is not much protection if the agent can issue twenty refunds of $199, so limits need to apply at more than one level.
The goal is simple: if the AI makes a mistake, limit how much damage that mistake can cause.
Control where it can send information. An agent that can read sensitive information should not automatically be able to send it anywhere it wants.
That does not stop someone from tricking the AI into following harmful instructions, or prevent every kind of attack. But limiting where information can go can reduce the damage if something goes wrong.
None of these ideas are new.
Security has worked this way for years: give a system only the access it needs, limit what it can do, and limit the damage when something fails.
The point is that these safeguards do not require a person to check every action one by one.
People still decide what the AI is allowed to access, how much it can spend, and what actions it can take on its own.
Human judgment is still there. It is used to set the boundaries in advance, rather than to approve every decision afterward.
Spend human review where it matters
None of this means human oversight is unnecessary.
Rules and limits can handle many situations, but they cannot cover every case.
A refund might be under the allowed amount and still look wrong. An action might follow every rule and still make no sense once you see the full situation. Something unexpected may happen that nobody thought to plan for.
That is where a person can still add value.
But if we want people to catch those unusual cases, we should not use up their attention reviewing hundreds of routine ones.
A reviewer checking every $40 refund has little attention left for the one $4,000 refund that deserves a careful look.
Human review works best when the decisions reaching a person are relatively rare, serious, and clear enough to evaluate.
So the design problem is not how to put a human in front of every action. It is how to shrink the number of actions that need one.
Let the AI operate inside tightly defined boundaries. Send the unusual or high-risk cases to a person. Give that person enough information and enough time to make a decision they can stand behind.
Security teams learned a similar lesson years ago. When analysts were faced with too many warnings, the solution was not to ask them to work harder. It was to reduce the noise and bring the most important warnings to their attention first.
AI agents need the same discipline.

Some things will still go wrong.
That is where a much larger argument happening in AI becomes relevant.
In September, Anthropic CEO Dario Amodei argued that improvements in the most capable AI systems may need to slow enough for safety work to keep up. He was explicit that slowing the pace does not mean stopping progress.
OpenAI CEO Sam Altman agreed that the development of the most capable systems may need to be paced. A few weeks later, though, he made the trade-off more explicit. In an interview reported by Reuters , he argued that society should accept some harm in exchange for the benefits of AI and people’s ability to use it.
The two differ on how much regulation and restraint that requires.
But underneath that disagreement is something useful for the much smaller problem here: neither is promising a world in which nothing goes wrong.
Some risk will remain.
The practical question is where we allow it.
Some failures can be handled through human review. Others should be made difficult by access restrictions, limits, isolation, or other safeguards before a person ever has to notice them.
Human review should sit where judgment adds something the system cannot easily provide—not where it is being used to compensate for an AI system that was given too much power in the first place.
What this does not solve
Limits can reduce what an AI system is capable of doing.
They do not explain why it behaved the way it did.
Records can help reconstruct what happened, but they may not give us a complete explanation of the model’s reasoning.
Human reviewers can catch mistakes.
They can also miss them.
There is no single safeguard that solves all of this.
That was the argument in my previous article too: AI safety depends on several layers, and those layers should not all fail for the same reason.
The part I would add now is that a human reviewer should not be treated as an unlimited safety layer.
The reviewer from the beginning of this article still has forty seconds and a $170 refund waiting on the screen.
Maybe the question is not simply whether a human approves it.
It is why this decision needed a human in the first place—and what would limit the damage if that human stopped looking.
Thanks to Wangden Bhutia, whose argument about agent communication outrunning human comprehension prompted this piece.