Figures are labelled MEASURED (observed data with a stated population), SURVEY (self-reported answers to a survey) or PREDICTION (an analyst forecast). Figures about AI agents directly are marked AGENT; the rest are labelled analogues from adjacent fields. A negative finding is a finding of search on the date given, not a statement that no such figure exists.
Summary
- The cost of step-by-step approval is measured in adjacent fields, and it is large. A 2006 review of 17 papers found drug safety alerts overridden in 49 to 96 per cent of cases (van der Sijs et al., JAMIA 2006); 63.77 per cent of 102,887 emergency-department medication alerts were overridden (JMIR Med Inform 2020). In software code review, Google reports a median wait for initial feedback of under an hour for small changes and about 5 hours for very large ones, and cites a median time to approval of 24 hours at Microsoft (Sadowski et al., ICSE-SEIP 2018).
- No measured human-time figures for agent approvals were located. A search of public framework documentation, vendor documentation and practitioner sources on 26 and 27 September 2026 located no published measured figure for approvals per 100 or 1,000 agent actions, median seconds to decision, or queue depth in production agent systems. A spot check on 2 October 2026 found recommended targets in practitioner guides, not measured production data.
- The human gate is expanding. In KPMG's Q1 2026 AI Quarterly Pulse (n=237 US leaders at firms with $1 billion or more in revenue), 63 per cent say they now require human validation of AI agent outputs, up from 22 per cent in Q1 2025 (KPMG, 31 Mar 2026). In Q2 2026 (n=204), 43 per cent name human oversight skills (for example, human-in-the-loop judgment and escalation skills) as a top challenge to deploying AI agents, second to data readiness and access at 58 per cent (KPMG Q2 2026). Gartner predicts over 40 per cent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls (Gartner, 25 Jun 2025).
- Human oversight does not reliably improve decision quality. In a controlled experiment (N=292), a human-in-the-loop design increased uptake of algorithmic recommendations but decreased decision accuracy, and monitors were less likely to adjust the least accurate recommendations (Sele and Chugunova, PLOS ONE 2024).
- Fewer, more specific alerts are accepted more often. In one vendor-published deployment, re-targeted hyperkalemia alerts cut alert volume 70 per cent and lowered the override rate from 81 to 38 per cent (FDB, 18 Apr 2023).
1. Approvals per unit of work
| Figure | Value / unit | Population and method | Date | Class | Source |
|---|---|---|---|---|---|
| AGENT: leaders requiring human validation of agent outputs | 63 per cent of respondents (up from 22 per cent in Q1 2025) | KPMG AI Quarterly Pulse Survey, Q1 2026; 237 US C-suite and business leaders at organisations with $1 billion or more in revenue; fielded 17 Feb to 17 Mar 2026 | 31 Mar 2026 | SURVEY | KPMG |
| AGENT: managing agent risk in the next 6 to 12 months by taking a "human-in-the-loop" approach where "a human validates outputs but does not oversee each agentic action or decision" | 57 per cent of respondents | KPMG Q1 2026 pulse, same panel | Mar 2026 | SURVEY | KPMG Q1 2026 |
| AGENT: human oversight skills named a top challenge to deploying AI agents | 43 per cent of respondents (behind data readiness and access at 58 per cent) | KPMG Q2 2026 pulse; 204 US leaders at organisations with $1 billion or more in revenue; fielded 28 Apr to 25 May 2026 | Jun 2026 | SURVEY | KPMG Q2 2026 |
| AGENT: approval events per 100 or 1,000 agent actions in production | No published measured figure located | Search of public framework, vendor and practitioner sources | 26 and 27 Sep 2026 | (negative finding) |
2. Time per approval (agent data absent; analogues labelled)
| Figure | Value / unit | Population and method | Date | Class | Source |
|---|---|---|---|---|---|
| Analogue, code review at Google: wait for initial feedback | Median under an hour for small changes; about 5 hours for very large changes; overall median latency for the entire review under 4 hours | Review logs for about 9 million changes, January 2014 to July 2016 | 2018 | MEASURED | Sadowski et al., ICSE-SEIP 2018 |
| Analogue, code review at Microsoft: time to approval | Median 24 hours (Czerwonka et al. 2015, as cited by Sadowski et al.); 14.7, 19.8 and 18.9 hours for three Microsoft projects (Rigby and Bird 2013, as cited) | Studies of Microsoft review data, cited in the Google paper | 2013 to 2015 | MEASURED | Sadowski et al., ICSE-SEIP 2018 |
| Analogue, code review: pull request pickup time benchmarks | Elite under 1 hour; good 1 to 4 hours; fair 5 to 16 hours; needs focus over 16 hours | LinearB 2026 benchmarks: more than 8.1 million pull requests from 4,813 teams and 163,820 contributors in 42 countries | 2026 | MEASURED (vendor benchmark) | LinearB |
| AGENT: median seconds from approval request to human decision in production agent systems | No published measured figure located | Search of public framework, vendor and practitioner sources | 26 and 27 Sep 2026 | (negative finding) |
3. Queue and wait time
| Figure | Value / unit | Population and method | Date | Class | Source |
|---|---|---|---|---|---|
| AGENT: how long a gated run can wait | Agent frameworks hold a paused run until a person responds. Microsoft's Copilot Studio team describes a human-in-the-loop workflow that "can wait for minutes, hours, or days"; LangGraph documentation states the "Graph waits indefinitely until you resume execution with a response". Distribution of actual waits in production: no published measured figure located (search of 26 and 27 Sep 2026) | Vendor documentation | 2026 | Capability as documented; no distribution published | Microsoft Copilot Studio CAT blog, 20 May 2026, LangGraph docs |
4. Approval fatigue and rubber-stamping
| Figure | Value / unit | Population and method | Date | Class | Source |
|---|---|---|---|---|---|
| Analogue, clinical: override rate | Drug safety alerts overridden in 49 to 96 per cent of cases | Systematic review, 17 papers, computerised physician order entry | 2006 | MEASURED (range across studies) | van der Sijs et al., JAMIA 2006 |
| Analogue, clinical: emergency department override | 63.77 per cent (65,616 of 102,887 alerts); 13.75 alerts per 100 orders | Retrospective descriptive study, tertiary urban academic emergency department, 18 months, 611 physicians, 71,546 patients | 2020 | MEASURED | JMIR Med Inform 2020 |
| Analogue, clinical: override by chart review | 92.9 per cent overridden (355 of 382); only 7.3 per cent of alerts clinically appropriate (28 of 382) | Retrospective observational study, stratified sampling, medical record review, inter-rater kappa reported | 2022 | MEASURED | Park et al., JMIR Med Inform 2022 |
| Analogue, clinical: drug allergy alerts over a decade | Override rate rose from 83.3 per cent (2004) to 87.6 per cent (2013); alerts for immune-mediated and life-threatening reactions with definite matches overridden 72.8 and 74.1 per cent of the time | 611,192 drug allergy alerts, two academic hospitals in Boston | 2016 | MEASURED | Topaz et al., JAMIA 2016 |
| Analogue, clinical: inappropriate dismissal | 73.3 per cent of patient allergy, drug-drug interaction and duplicate drug alerts overridden; about 40 per cent of overrides not appropriate | Three-year study, 793-bed tertiary-care teaching institution | 2018 | MEASURED | Nanji et al., JAMIA 2018 |
| Analogue, MFA: push-fatigue approvals | Under repeated push prompts, "the user approves a request, sometimes thinking it's a legitimate login attempt, or just to stop the noise"; no controlled effect size published in the sources read. Microsoft enforced number matching for all Authenticator push notifications from 8 May 2023 in response to MFA fatigue attacks | Vendor security literature and reporting | 2023 to 2025 | Mechanism described; vendor response | WorkOS, 9 Oct 2025, BleepingComputer, 8 May 2023 |
| Analogue, code review at volume: review quality falls with size | A SmartBear study of a Cisco Systems programming team found developers should review no more than 200 to 400 lines of code at a time; "a review of 200-400 LOC over 60 to 90 minutes should yield 70-90% defect discovery" | Vendor-published case study | Undated page | MEASURED (vendor-published) | SmartBear |
| AGENT-adjacent: human monitors under-adjust the worst errors | Monitors adjusted 64 per cent of recommendations with a small error and 60 per cent with a large error (p=0.02); adjustments were 11.3 percentile points for smaller errors vs 9.6 for larger ones (p<0.001); overall accuracy worse with human-in-the-loop than with pure delegation (mean absolute deviation 18.0 vs 17.4, p=0.04) | Controlled online experiment, N=292, between-subjects | 2024 | MEASURED (experiment) | Sele and Chugunova, PLOS ONE 2024 |
5. Abandonment
| Figure | Value / unit | Population and method | Date | Class | Source |
|---|---|---|---|---|---|
| AGENT: projected cancellations | Over 40 per cent of agentic AI projects cancelled by the end of 2027, citing escalating costs, unclear business value or inadequate risk controls | Gartner prediction; supporting January 2025 poll of 3,412 webinar attendees: 19 per cent significant investment in agentic AI, 42 per cent conservative | 25 Jun 2025 | PREDICTION | Gartner |
| AGENT: measured abandonment because approval volume was unworkable | No published measured dataset located. Practitioner sources describe rubber-stamping risk and teams narrowing approval steps, but no measured dataset of teams removing the human step was found | Search of public sources | 26 and 27 Sep 2026 | (negative finding) | |
| Analogue: alert-volume reduction | 70 per cent fewer alerts (15,057 traditional alerts to 4,590 targeted alerts over six months); override rate fell from 81 to 38 per cent; acceptance rose from under 20 to over 60 per cent | Targeted hyperkalemia alerts, Lehigh Valley Health Network | 18 Apr 2023 | MEASURED (vendor-published deployment data) | FDB |
6. The counter-case
- Fewer, more specific alerts. In the FDB deployment, when alerts were fewer and more specific, acceptance rose from under 20 to over 60 per cent and alert volume fell 70 per cent. MEASURED, vendor-published (FDB, 2023).
- Buyers want the gate. 63 per cent of large-enterprise leaders say they require human validation of agent outputs, up from 22 per cent a year earlier. SURVEY (KPMG Q1 2026). The demand for the human step is growing while its cost in agent systems is unmeasured in public data.
- Human-in-the-loop increases adoption. In the N=292 experiment, offering monitor-and-adjust control raised the preference for the algorithm by 7 percentage points (66 to 73 per cent) and raised user confidence (48 vs 41 out of 100, p=0.008): oversight bought trust even where it did not buy accuracy. MEASURED (Sele and Chugunova 2024).
- The quality caveat is equally strong evidence. The same experiment found human-in-the-loop reduced accuracy and that monitors adjusted the largest errors less often. The controlled evidence located does not support a claim that step-by-step approval improves outcomes.
7. Adoption data (all SURVEY, populations stated)
| Survey | Figure | Population | Date | Source |
|---|---|---|---|---|
| McKinsey, The state of AI in 2025 | 62 per cent at least experimenting with AI agents; 23 per cent scaling an agentic AI system somewhere in their enterprises; 88 per cent report regular AI use in at least one business function | 1,993 participants in 105 nations; fielded 25 Jun to 29 Jul 2025 | Nov 2025 | McKinsey |
| KPMG AI Pulse Q1 2026 | 63 per cent require human validation of agent outputs (from 22 per cent in Q1 2025); 57 per cent taking a "human-in-the-loop" approach where a human validates outputs | 237 US leaders at $1 billion+ firms | Mar 2026 | KPMG |
| KPMG AI Pulse Q2 2026 | AI agent deployment 53 per cent (55 per cent in Q1); scaling AI agents across multiple functions 29 per cent (33 per cent in Q1); orchestrating multiple agents across workflows 18 per cent (9 per cent in Q1); piloting 35 per cent (30 per cent in Q1) | 204 US leaders at $1 billion+ firms | Jun 2026 | KPMG Q2 2026 |
| Gartner | 17 per cent of organisations have deployed AI agents (2026 Gartner CIO and Technology Executive Survey, 2,501 respondents, fielded May to June 2025). PREDICTION: at least 15 per cent of day-to-day work decisions made autonomously through agentic AI by 2028; Gartner estimates only about 130 of thousands of vendors claiming agentic AI are genuinely agentic | CIO survey n=2,501; webinar poll n=3,412 | 2025 to 2026 | Gartner, Hype Cycle for Agentic AI, Gartner, 25 Jun 2025 |
| Deloitte, State of AI in the Enterprise | 25 per cent have moved 40 per cent or more of their AI pilots into production | 3,235 business and IT leaders in 24 countries; fielded Aug to Sep 2025 | 21 Jan 2026 | Deloitte |
| PwC AI Agent Survey | 79 per cent say AI agents are already being adopted in their companies | 308 US executives; fielded 22 to 28 Apr 2025 | 2025 | PwC |
The human-validation figure in KPMG's survey rose from 22 to 63 per cent in four quarters. The surveys define adoption and deployment differently, so their headline figures are not directly comparable.
8. Standards and regulators
| Body / instrument | Date | What it says on human control | Source |
|---|---|---|---|
| NIST Center for AI Standards and Innovation (CAISI), AI Agent Standards Initiative | Announced 17 Feb 2026; RFI on AI agent security closed 9 Mar 2026; concept paper on AI agent identity and authorization comments closed 2 Apr 2026 | Three pillars: industry-led agent standards and US leadership in international standards bodies; community-led open-source protocol development; research on AI agent security and identity | NIST |
| NIST technical blog, AI agent hijacking evaluations | 17 Jan 2025 | Red-team attacks raised hijacking success "from 11% for the strongest baseline attack to 81% for the strongest new attack" against an agent built on the upgraded Claude 3.5 Sonnet. MEASURED | NIST |
| OWASP Agent Control Standard (ACS) | 1 Sep 2026 | Agents must be "inspectable, traceable and instrumentable"; the standard enables "declarative controls that are portable across agent frameworks and enforced at runtime" | OWASP GenAI Security Project |
| EU AI Act, Articles 12, 14, 26 | In force since 1 Aug 2024. As amended by the Digital Omnibus in 2026, Annex III high-risk obligations apply from 2 Dec 2027 and Annex I (regulated products) from 2 Aug 2028, as reported | Art. 14(1): high-risk systems "shall be designed and developed in such a way, including with appropriate human-machine interface tools, that they can be effectively overseen by natural persons during the period in which they are in use." Art. 14(4)(b) names automation bias. Art. 14(4)(e): overseers must be able "to intervene in the operation of the high-risk AI system or interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state." Art. 14(5): for remote biometric identification systems, no action or decision on an identification unless it is separately verified and confirmed by at least two natural persons. Art. 12(1): high-risk systems "shall technically allow for the automatic recording of events (logs) over the lifetime of the system." Art. 26(6): deployers keep the automatically generated logs, to the extent they are under their control, for at least six months | Art. 14, Art. 12, Art. 26, Orrick on the Digital Omnibus, Jul 2026 |
| Draft EU GMP Annex 22, Artificial Intelligence | Published for consultation 7 Jul 2025; consultation closed 7 Oct 2025 | Applies to static models with deterministic output in critical GMP applications; dynamic models, probabilistic-output models, generative AI and large language models "should not be used in critical GMP applications". In non-critical use of generative AI and LLMs, "personnel with adequate qualification and training should always be responsible for ensuring that the outputs from such models are suitable for the intended use". Where a model gives input to a human operator's decision and testing effort has been diminished, records should be kept and, depending on criticality, "this may imply a consistent review and/or test of every output from the model". A 4-eyes principle applies where staff with access to test data must also work on training and validation | European Commission, draft Annex 22 (PDF), consultation page |
| 21 CFR Part 11 | Current | Secure, computer-generated, time-stamped audit trails (11.10(e)); signed records show the signer's name, the date and time, and the meaning of the signature (11.50); signatures linked to their records so they cannot be excised, copied or transferred (11.70) | 11.10, 11.50, 11.70 |
| FDA draft guidance on AI to support regulatory decision-making for drugs and biologics | Jan 2025, draft, non-binding | Proposes "a risk-based credibility assessment framework" for establishing and evaluating the credibility of an AI model for a particular context of use | FDA |
Figures labelled SURVEY or PREDICTION should not be quoted without the label attached. Negative findings are findings of search on the dates given, not assertions of non-existence.
