I wrote earlier about why the human half of that equation is not one variable, since availability, competence, and state all fail independently. This article goes further and asks what happens once that variable human and an imperfect AI are actually working together, in an enterprise workflow or a government process. The evidence is specific: oversight sometimes closes real gaps and sometimes makes outcomes worse than either party would have produced alone, and "a human is present" tells you almost nothing about which one you're getting.

For this article, oversight means a final-stage checkpoint: a human reviewing a completed AI output before it ships or is acted on, not real-time monitoring or post-hoc auditing, which carry different failure mechanisms.

The evidence below is strongest in the domains it was studied in: high-stakes, non-routine judgments where a miss has real cost, such as diagnosis, financial decisions, safety-critical code, and anything a regulator would call high-risk. Nobody needs a rigorous review protocol for a drafted email or a meeting summary, and low-cost theater is a defensible trade-off there. The limit on that defense is that the habit doesn't respect the boundary: a reviewer who rubber-stamps routine summaries all week doesn't switch into a differently calibrated reviewer when a high-stakes case lands in the same queue.

The EU AI Act's Article 14 requires human oversight for high-risk systems, and for acute applications like remote biometric identification it requires independent verification by at least two people, not one. NIST's AI Risk Management Framework recommends oversight more generally, and most enterprise AI governance policies I have seen in the past year include some version of a human checkpoint before a high-stakes output ships. The clearest test of whether that checkpoint delivers what it promises is a 2024 meta-analysis in Nature Human Behaviour, covering 106 experimental studies and 370 effect sizes on human-AI collaboration. The headline result: human-AI combinations performed significantly worse on average than the best of the human or the AI working alone (Hedges' g = -0.23, 95% CI -0.39 to -0.07).

That comparison deserves a caveat. Oversight policy isn't really trying to beat a "best of either party" benchmark; in production nobody knows in advance which party is stronger on a given case, which is exactly why a checkpoint exists. The more useful part of the meta-analysis is what moderated the effect. Task type mattered: decision-making tasks showed real losses from combination, content-creation tasks showed real gains. Relative skill mattered more. When the human was independently the stronger party, collaboration produced gains. When the AI was independently the stronger party, collaboration produced losses. That doesn't remove the uncertainty about which party is stronger in a given case, but it tells you the answer to that question, not the mere presence of a reviewer, is what predicts whether the checkpoint helps.

Where it fails in practice

Three of the Big Four professional services firms and three separate government bodies have made this case for me in the past ten months. Deloitte, EY, and KPMG each had reports withdrawn or refunded after fabricated citations surfaced. The EU's cybersecurity agency, ENISA, admitted two of its 2025 threat reports contained AI-hallucinated sources. South Africa withdrew two separate cabinet-approved policy documents within a week of each other over the same failure, one of them with a higher fabrication rate than the one before it. Every one of these organisations had a defined review process on paper. None of it caught the error before publication. The question worth asking is not whether that keeps happening, since it evidently does, but why a process that exists so reliably fails to function.

The checkpoint that didn't catch itself In June 2026, KPMG pulled a flagship report on agentic AI adoption after a forensic audit found only five of its 45 citations pointed to real, intact sources. A KPMG spokesperson responded that the firm requires staff to follow mandatory human oversight to verify content and cross-check sources. That policy did not catch the error in the firm's own report.

The failure mode has a name in the human factors literature: automation bias, the tendency to defer to an automated recommendation and reduce independent scrutiny accordingly. A named human reviewer also does a second job that has nothing to do with catching errors: for a board or regulator, they're a place for liability to land if something goes wrong. A checkpoint can serve that function well while serving its stated safety function poorly, and the two get conflated more often than they should.

A 2026 preprint, not yet peer reviewed, studied developers monitoring AI coding agents for sabotage over realistic five-hour tasks. Baseline: 94 percent failed to detect the sabotage at all. When researchers added a safety monitor that correctly flagged the malicious action, 56 percent of participants still accepted the flagged code anyway. The checkpoint worked as designed, and the human overrode it in the majority of cases, largely because the monitor delivered a passive alert rather than a forcing function, requiring nothing different of the reviewer before they proceeded, which is exactly the distinction the friction research below identifies as decisive. I'm treating this as an early signal given its review status, but it lines up with older, peer-reviewed automation-bias research going back over a decade.

And the failure is not limited to inattention. A double-reading mammography study, published in Cancer Imaging, randomized whether a second radiologist was told the first read came from a colleague or from an AI system, and found that disclosure changed how often the second reader overturned an incorrect first read. Knowing a machine had already looked measurably changed how hard the human looked next. This happens to careful, credentialed professionals under review conditions specifically designed to catch errors.

The three conditions that separate working oversight from theater

Set against that record, the cases where human-AI collaboration clearly outperforms either party alone share three characteristics, not one.

Relative skill on the specific task, not general qualification

It's worth distinguishing this from a more forgiving standard: whether a human-AI pairing beats the human alone (augmentation) is a much lower bar than whether it beats the AI alone (synergy), and oversight requirements generally assume the latter while often only achieving the former. A 2026 study in Radiology makes the distinction concrete. Researchers tested an AI-assisted mammography workflow on 60 general radiologists and 35 breast imaging specialists. Among general radiologists, the AI workflow raised cancer detection from 3.76 to 4.99 per 1,000 exams, closing, and slightly exceeding, the gap to specialists' own rate of 4.76 per 1,000. Among the specialists, it produced no significant change at all. The AI didn't make good readers better; it brought weaker readers up to where the stronger ones already were. A single workflow, applied uniformly to 95 radiologists, produced a real gain for 60 of them and nothing measurable for the other 35. The mandate was uniform. The effect was not.

A related, separately-run mammography study reinforces the point from the other direction: using AI as a second reader alongside a human first reader improved sensitivity from 87.4 to 91.8 percent over human-human double reading, with the largest gains among less experienced radiologists. Specificity decreased slightly, from 84.0 to 83.0 percent, not statistically significant in that study, but worth naming: oversight that catches more of what the AI missed tends to also flag more of what the AI got right. Any account of an oversight benefit that reports only what it caught, not what it cost, is telling half the story.

What the reviewer is shown

Research on Amplified Oversight, including DeepMind-affiliated authors, found that assistance stripping away the AI's conclusions and preserving only its evidence-gathering was uniquely effective at improving human judgment, while assistance including the AI's reasoning, ratings, and stated confidence measurably decreased human accuracy below an unassisted baseline. In practice, evidence-gathering assistance looks like surfacing the underlying source passages or prior cases a system consulted and leaving the reviewer to draw their own conclusion, rather than surfacing the system's answer with a confidence score attached. Showing a reviewer what a system found tends to help. Showing a reviewer what a system concluded tends to replace the reviewer's own conclusion, which is the opposite of oversight.

Friction the reviewer actually experiences as friction

It needs to pay the reviewer back with a calibration signal, not just slow them down. A 2021 CSCW study on cognitive forcing functions, brief mandatory justification steps or delayed reveal of the AI's suggestion, found they significantly reduced overreliance compared to standard explainable-AI approaches. The same study found participants rated these designs least favorably of all the conditions tested, which is the real design problem: friction a reviewer resents gets bypassed once volume pressure sets in, and a queue cleared by fast rubber-stamping is worse than no checkpoint at all, because it produces an audit trail that looks like review without the substance of it.

What this means

Human review still matters where the stakes are real, and nothing above argues for cutting it. What the evidence supports is a stricter test for the phrase itself: "a human is in the loop" is not, by itself, a claim about safety. It becomes true or false only once it specifies which human, reviewing what evidence, under what friction, on a task where that human has a real edge. A checkpoint that shows a reviewer a confident AI conclusion, staffed by someone with no particular advantage on that judgment, under time pressure that discourages disagreement, is not a weaker version of oversight. It is closer to no oversight at all, dressed as compliance, whether it sits inside an enterprise workflow or a government process. The stakes differ, a rubber-stamped queue produces a bad financial decision in business and a wrongly flagged or denied person under a safeguard like the EU AI Act's biometric identification rule, but the mechanism is the same failure wearing a different uniform: a human present in name, without the conditions that make their judgment count.

The presence of a person in the workflow has never been the variable that mattered. What mattered, in every study and every incident above, was something more specific underneath it, and regulation, procurement, and governance policy have mostly been written as though presence were the whole answer.