Skip to content
Regulayer™Human Control for AI
Book a live demo

Research library · Human oversight

The Human Cost of Approving AI Agent Actions

Public evidence, gathered on 26 and 27 September 2026 and checked on 2 October 2026, on what step-by-step human approval costs: alert override and rubber-stamping in clinical decision support, review queues in software code review, push-approval fatigue, a controlled experiment on human-in-the-loop accuracy, enterprise surveys on human validation of AI agent outputs, and what standards and regulators say about human control.

Compiled from public sources, 27 September 2026. Information, not legal advice.

Figures are labelled MEASURED (observed data with a stated population), SURVEY (self-reported answers to a survey) or PREDICTION (an analyst forecast). Figures about AI agents directly are marked AGENT; the rest are labelled analogues from adjacent fields. A negative finding is a finding of search on the date given, not a statement that no such figure exists.

Summary

  • The cost of step-by-step approval is measured in adjacent fields, and it is large. A 2006 review of 17 papers found drug safety alerts overridden in 49 to 96 per cent of cases (van der Sijs et al., JAMIA 2006); 63.77 per cent of 102,887 emergency-department medication alerts were overridden (JMIR Med Inform 2020). In software code review, Google reports a median wait for initial feedback of under an hour for small changes and about 5 hours for very large ones, and cites a median time to approval of 24 hours at Microsoft (Sadowski et al., ICSE-SEIP 2018).
  • No measured human-time figures for agent approvals were located. A search of public framework documentation, vendor documentation and practitioner sources on 26 and 27 September 2026 located no published measured figure for approvals per 100 or 1,000 agent actions, median seconds to decision, or queue depth in production agent systems. A spot check on 2 October 2026 found recommended targets in practitioner guides, not measured production data.
  • The human gate is expanding. In KPMG's Q1 2026 AI Quarterly Pulse (n=237 US leaders at firms with $1 billion or more in revenue), 63 per cent say they now require human validation of AI agent outputs, up from 22 per cent in Q1 2025 (KPMG, 31 Mar 2026). In Q2 2026 (n=204), 43 per cent name human oversight skills (for example, human-in-the-loop judgment and escalation skills) as a top challenge to deploying AI agents, second to data readiness and access at 58 per cent (KPMG Q2 2026). Gartner predicts over 40 per cent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls (Gartner, 25 Jun 2025).
  • Human oversight does not reliably improve decision quality. In a controlled experiment (N=292), a human-in-the-loop design increased uptake of algorithmic recommendations but decreased decision accuracy, and monitors were less likely to adjust the least accurate recommendations (Sele and Chugunova, PLOS ONE 2024).
  • Fewer, more specific alerts are accepted more often. In one vendor-published deployment, re-targeted hyperkalemia alerts cut alert volume 70 per cent and lowered the override rate from 81 to 38 per cent (FDB, 18 Apr 2023).

1. Approvals per unit of work

Figure Value / unit Population and method Date Class Source
AGENT: leaders requiring human validation of agent outputs 63 per cent of respondents (up from 22 per cent in Q1 2025) KPMG AI Quarterly Pulse Survey, Q1 2026; 237 US C-suite and business leaders at organisations with $1 billion or more in revenue; fielded 17 Feb to 17 Mar 2026 31 Mar 2026 SURVEY KPMG
AGENT: managing agent risk in the next 6 to 12 months by taking a "human-in-the-loop" approach where "a human validates outputs but does not oversee each agentic action or decision" 57 per cent of respondents KPMG Q1 2026 pulse, same panel Mar 2026 SURVEY KPMG Q1 2026
AGENT: human oversight skills named a top challenge to deploying AI agents 43 per cent of respondents (behind data readiness and access at 58 per cent) KPMG Q2 2026 pulse; 204 US leaders at organisations with $1 billion or more in revenue; fielded 28 Apr to 25 May 2026 Jun 2026 SURVEY KPMG Q2 2026
AGENT: approval events per 100 or 1,000 agent actions in production No published measured figure located Search of public framework, vendor and practitioner sources 26 and 27 Sep 2026 (negative finding)

2. Time per approval (agent data absent; analogues labelled)

Figure Value / unit Population and method Date Class Source
Analogue, code review at Google: wait for initial feedback Median under an hour for small changes; about 5 hours for very large changes; overall median latency for the entire review under 4 hours Review logs for about 9 million changes, January 2014 to July 2016 2018 MEASURED Sadowski et al., ICSE-SEIP 2018
Analogue, code review at Microsoft: time to approval Median 24 hours (Czerwonka et al. 2015, as cited by Sadowski et al.); 14.7, 19.8 and 18.9 hours for three Microsoft projects (Rigby and Bird 2013, as cited) Studies of Microsoft review data, cited in the Google paper 2013 to 2015 MEASURED Sadowski et al., ICSE-SEIP 2018
Analogue, code review: pull request pickup time benchmarks Elite under 1 hour; good 1 to 4 hours; fair 5 to 16 hours; needs focus over 16 hours LinearB 2026 benchmarks: more than 8.1 million pull requests from 4,813 teams and 163,820 contributors in 42 countries 2026 MEASURED (vendor benchmark) LinearB
AGENT: median seconds from approval request to human decision in production agent systems No published measured figure located Search of public framework, vendor and practitioner sources 26 and 27 Sep 2026 (negative finding)

3. Queue and wait time

Figure Value / unit Population and method Date Class Source
AGENT: how long a gated run can wait Agent frameworks hold a paused run until a person responds. Microsoft's Copilot Studio team describes a human-in-the-loop workflow that "can wait for minutes, hours, or days"; LangGraph documentation states the "Graph waits indefinitely until you resume execution with a response". Distribution of actual waits in production: no published measured figure located (search of 26 and 27 Sep 2026) Vendor documentation 2026 Capability as documented; no distribution published Microsoft Copilot Studio CAT blog, 20 May 2026, LangGraph docs

4. Approval fatigue and rubber-stamping

Figure Value / unit Population and method Date Class Source
Analogue, clinical: override rate Drug safety alerts overridden in 49 to 96 per cent of cases Systematic review, 17 papers, computerised physician order entry 2006 MEASURED (range across studies) van der Sijs et al., JAMIA 2006
Analogue, clinical: emergency department override 63.77 per cent (65,616 of 102,887 alerts); 13.75 alerts per 100 orders Retrospective descriptive study, tertiary urban academic emergency department, 18 months, 611 physicians, 71,546 patients 2020 MEASURED JMIR Med Inform 2020
Analogue, clinical: override by chart review 92.9 per cent overridden (355 of 382); only 7.3 per cent of alerts clinically appropriate (28 of 382) Retrospective observational study, stratified sampling, medical record review, inter-rater kappa reported 2022 MEASURED Park et al., JMIR Med Inform 2022
Analogue, clinical: drug allergy alerts over a decade Override rate rose from 83.3 per cent (2004) to 87.6 per cent (2013); alerts for immune-mediated and life-threatening reactions with definite matches overridden 72.8 and 74.1 per cent of the time 611,192 drug allergy alerts, two academic hospitals in Boston 2016 MEASURED Topaz et al., JAMIA 2016
Analogue, clinical: inappropriate dismissal 73.3 per cent of patient allergy, drug-drug interaction and duplicate drug alerts overridden; about 40 per cent of overrides not appropriate Three-year study, 793-bed tertiary-care teaching institution 2018 MEASURED Nanji et al., JAMIA 2018
Analogue, MFA: push-fatigue approvals Under repeated push prompts, "the user approves a request, sometimes thinking it's a legitimate login attempt, or just to stop the noise"; no controlled effect size published in the sources read. Microsoft enforced number matching for all Authenticator push notifications from 8 May 2023 in response to MFA fatigue attacks Vendor security literature and reporting 2023 to 2025 Mechanism described; vendor response WorkOS, 9 Oct 2025, BleepingComputer, 8 May 2023
Analogue, code review at volume: review quality falls with size A SmartBear study of a Cisco Systems programming team found developers should review no more than 200 to 400 lines of code at a time; "a review of 200-400 LOC over 60 to 90 minutes should yield 70-90% defect discovery" Vendor-published case study Undated page MEASURED (vendor-published) SmartBear
AGENT-adjacent: human monitors under-adjust the worst errors Monitors adjusted 64 per cent of recommendations with a small error and 60 per cent with a large error (p=0.02); adjustments were 11.3 percentile points for smaller errors vs 9.6 for larger ones (p<0.001); overall accuracy worse with human-in-the-loop than with pure delegation (mean absolute deviation 18.0 vs 17.4, p=0.04) Controlled online experiment, N=292, between-subjects 2024 MEASURED (experiment) Sele and Chugunova, PLOS ONE 2024

5. Abandonment

Figure Value / unit Population and method Date Class Source
AGENT: projected cancellations Over 40 per cent of agentic AI projects cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls Gartner prediction; supporting January 2025 poll of 3,412 webinar attendees: 19 per cent significant investment in agentic AI, 42 per cent conservative 25 Jun 2025 PREDICTION Gartner
AGENT: measured abandonment because approval volume was unworkable No published measured dataset located. Practitioner sources describe rubber-stamping risk and teams narrowing approval steps, but no measured dataset of teams removing the human step was found Search of public sources 26 and 27 Sep 2026 (negative finding)
Analogue: alert-volume reduction 70 per cent fewer alerts (15,057 traditional alerts to 4,590 targeted alerts over six months); override rate fell from 81 to 38 per cent; acceptance rose from under 20 to over 60 per cent Targeted hyperkalemia alerts, Lehigh Valley Health Network 18 Apr 2023 MEASURED (vendor-published deployment data) FDB

6. The counter-case

  1. Fewer, more specific alerts. In the FDB deployment, when alerts were fewer and more specific, acceptance rose from under 20 to over 60 per cent and alert volume fell 70 per cent. MEASURED, vendor-published (FDB, 2023).
  2. Buyers want the gate. 63 per cent of large-enterprise leaders say they require human validation of agent outputs, up from 22 per cent a year earlier. SURVEY (KPMG Q1 2026). The demand for the human step is growing while its cost in agent systems is unmeasured in public data.
  3. Human-in-the-loop increases adoption. In the N=292 experiment, offering monitor-and-adjust control raised the preference for the algorithm by 7 percentage points (66 to 73 per cent) and raised user confidence (48 vs 41 out of 100, p=0.008): oversight bought trust even where it did not buy accuracy. MEASURED (Sele and Chugunova 2024).
  4. The quality caveat is equally strong evidence. The same experiment found human-in-the-loop reduced accuracy and that monitors adjusted the largest errors less often. The controlled evidence located does not support a claim that step-by-step approval improves outcomes.

7. Adoption data (all SURVEY, populations stated)

Survey Figure Population Date Source
McKinsey, The state of AI in 2025 62 per cent at least experimenting with AI agents; 23 per cent scaling an agentic AI system somewhere in their enterprises; 88 per cent report regular AI use in at least one business function 1,993 participants in 105 nations; fielded 25 Jun to 29 Jul 2025 Nov 2025 McKinsey
KPMG AI Pulse Q1 2026 63 per cent require human validation of agent outputs (from 22 per cent in Q1 2025); 57 per cent taking a "human-in-the-loop" approach where a human validates outputs 237 US leaders at $1 billion+ firms Mar 2026 KPMG
KPMG AI Pulse Q2 2026 AI agent deployment 53 per cent (55 per cent in Q1); scaling AI agents across multiple functions 29 per cent (33 per cent in Q1); orchestrating multiple agents across workflows 18 per cent (9 per cent in Q1); piloting 35 per cent (30 per cent in Q1) 204 US leaders at $1 billion+ firms Jun 2026 KPMG Q2 2026
Gartner 17 per cent of organisations have deployed AI agents (2026 Gartner CIO and Technology Executive Survey, 2,501 respondents, fielded May to June 2025). PREDICTION: at least 15 per cent of day-to-day work decisions made autonomously through agentic AI by 2028; Gartner estimates only about 130 of thousands of vendors claiming agentic AI are genuinely agentic CIO survey n=2,501; webinar poll n=3,412 2025 to 2026 Gartner, Hype Cycle for Agentic AI, Gartner, 25 Jun 2025
Deloitte, State of AI in the Enterprise 25 per cent have moved 40 per cent or more of their AI pilots into production 3,235 business and IT leaders in 24 countries; fielded Aug to Sep 2025 21 Jan 2026 Deloitte
PwC AI Agent Survey 79 per cent say AI agents are already being adopted in their companies 308 US executives; fielded 22 to 28 Apr 2025 2025 PwC

The human-validation figure in KPMG's survey rose from 22 to 63 per cent in four quarters. The surveys define adoption and deployment differently, so their headline figures are not directly comparable.

8. Standards and regulators

Body / instrument Date What it says on human control Source
NIST Center for AI Standards and Innovation (CAISI), AI Agent Standards Initiative Announced 17 Feb 2026; RFI on AI agent security closed 9 Mar 2026; concept paper on AI agent identity and authorization comments closed 2 Apr 2026 Three pillars: industry-led agent standards and US leadership in international standards bodies; community-led open-source protocol development; research on AI agent security and identity NIST
NIST technical blog, AI agent hijacking evaluations 17 Jan 2025 Red-team attacks raised hijacking success "from 11% for the strongest baseline attack to 81% for the strongest new attack" against an agent built on the upgraded Claude 3.5 Sonnet. MEASURED NIST
OWASP Agent Control Standard (ACS) 1 Sep 2026 Agents must be "inspectable, traceable and instrumentable"; the standard enables "declarative controls that are portable across agent frameworks and enforced at runtime" OWASP GenAI Security Project
EU AI Act, Articles 12, 14, 26 In force since 1 Aug 2024. As amended by the Digital Omnibus in 2026, Annex III high-risk obligations apply from 2 Dec 2027 and Annex I (regulated products) from 2 Aug 2028, as reported Art. 14(1): high-risk systems "shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use." Art. 14(4)(b) names automation bias. Art. 14(4)(e): overseers must be able "to intervene in the operation of the high-risk AI system or interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state." Art. 14(5): for remote biometric identification systems, no action or decision on an identification unless it is separately verified and confirmed by at least two natural persons. Art. 12(1): high-risk systems "shall technically allow for the automatic recording of events (logs) over the lifetime of the system." Art. 26(6): deployers keep the automatically generated logs, to the extent they are under their control, for at least six months Art. 14, Art. 12, Art. 26, Orrick on the Digital Omnibus, Jul 2026
Draft EU GMP Annex 22, Artificial Intelligence Published for consultation 7 Jul 2025; consultation closed 7 Oct 2025 Applies to static models with deterministic output in critical GMP applications; dynamic models, probabilistic-output models, generative AI and large language models "should not be used in critical GMP applications". In non-critical use of generative AI and LLMs, "personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use". Where a model gives input to a human operator's decision and testing effort has been diminished, records should be kept and, depending on criticality, "this may imply a consistent review and/or test of every output from the model". A 4-eyes principle applies where staff with access to test data must also work on training and validation European Commission, draft Annex 22 (PDF), consultation page
21 CFR Part 11 Current Secure, computer-generated, time-stamped audit trails (11.10(e)); signed records show the signer's name, the date and time, and the meaning of the signature (11.50); signatures linked to their records so they cannot be excised, copied or transferred (11.70) 11.10, 11.50, 11.70
FDA draft guidance on AI to support regulatory decision-making for drugs and biologics Jan 2025, draft, non-binding Proposes "a risk-based credibility assessment framework" for establishing and evaluating the credibility of an AI model for a particular context of use FDA

Figures labelled SURVEY or PREDICTION should not be quoted without the label attached. Negative findings are findings of search on the dates given, not assertions of non-existence.