Human review only protects you if it sits where an agent's confidence and competence diverge, before output is polished into fluent prose. Batch approvals and late reviews are oversight in name only.
This closes the cluster this series has spent four pieces building: how an agent gets compromised, how far the damage can spread, and what it inherits inside a logged-in session. Every one of those pieces has pointed towards the same safeguard, a human somewhere in the loop, without ever asking the harder question this piece takes on directly: where, specifically, does that human actually add value, and where is their presence in the loop closer to a formality than a genuine check.
"Human-in-the-loop" gets used as though it were a single design choice, applied uniformly. In practice it is a placement decision, made separately for every task, and the placement is what determines whether the human is doing real work or just adding a click between the agent and the outcome.
Research on professionals using AI on tasks inside and outside the tool's actual competence found something specific and uncomfortable: on tasks outside that competence, people using the tool were 19 percentage points less likely to reach the correct answer than people not using it at all. The tool did not just fail to help there. It actively made outside-competence tasks worse, because people trusted output that looked as confident and well-formatted as the output on tasks the tool was actually good at. A human in the loop who cannot tell which kind of task they are looking at is not a safeguard. They are a rubber stamp with good intentions.
The human's judgement is worth the most at the exact point where the agent's confidence and its actual competence are most likely to diverge, which is rarely the point where the interface puts the approval button by default. A human reviewing a fully drafted, fluent, well-structured output is reviewing something engineered to look correct, whether or not it is. The same human, shown the underlying decision before the agent smoothed it into prose, catches far more, because the thing they are evaluating has not yet been dressed up to look finished.
A human in the loop is only a safeguard if the human is looking at the thing that is actually likely to be wrong, at the point before it has been polished into looking right. Everywhere else, it is a delay with the appearance of oversight.
Place human review at the point where an agent's confidence and its actual competence are most likely to diverge, before the output has been polished into fluent prose, rather than defaulting to whichever approval point the interface makes easiest to add.
A browser agent in a logged-in session inherits everything that login can reach. Judge the session, not the task, and default to no where irreversible or highly sensitive actions are available.
Custom Agents & Tools · 4 minFramework · 2 October 2026Before connecting an agent to a new system, ask what it could reach on its worst day. Grant access per task rather than per agent, and separate read, write and send, so one mistake stays contained.
Custom Agents & Tools · 4 minExplainer · 1 October 2026To an agent that reads it, any shared drive or ticket system many people could edit is untrusted content. Scope read access, curate trusted sources and log what the agent reads.
Custom Agents & Tools · 4 minIf this is the question on your desk, a thirty-minute call tells you whether the service fits, or that you do not need us yet.