Skip to content
Regulayer™Human Control for AI
Book a live demo

Research library · Human oversight

Human operator state and decision quality in AI-assisted regulated environments

A review of published evidence on how fatigue, cognitive load and stress affect error rates and oversight quality in clinical and other regulated settings, including automation bias when decision support is wrong. It also covers the relevant regulatory texts and existing rules that restrict duty on grounds of alcohol, fatigue or self-declared unfitness.

Compiled from public sources, February 2026. Information, not legal advice.

Scope and review approach

This review synthesizes published evidence on how human cognitive and physiological state (fatigue, cognitive load and stress) affects decision quality, error rates, and oversight effectiveness in clinical, pharmaceutical, and regulated laboratory-adjacent settings, with explicit attention to AI- and automation-assisted workflows and human-in-the-loop supervisory control. It also surveys regulatory and standards frameworks relevant to GxP-context deployment and identifies operational precedents where physiological state is used to restrict or gate high-stakes actions. [1]

In this review, the evidence base is strongest in (a) healthcare shift work and resident duty hours, where error effects are quantified, and (b) transportation, aviation and nuclear human-factors literature, where supervisory control and fatigue risk management are mature and often quantified. Direct, peer-reviewed evidence in robotic and automated laboratory operations and pharmaceutical manufacturing sign-offs exists but is comparatively sparse; where that gap exists, this review explicitly flags it and uses closest-domain analogs (robotic surgery, process control, aviation, nuclear operations).

Regulatory source material covered includes guidance from the U.S. Food and Drug Administration and the European Medicines Agency on "Good AI Practice in Drug Development," the EU Artificial Intelligence Act (notably Articles 9 and 14), and quality-system frameworks that frequently anchor GxP governance arguments, including International Council for Harmonisation Q10 and International Organization for Standardization ISO 13485.

Fatigue and error rates

A highly cited randomized, prospective intensive-care study quantified substantial error inflation when clinicians worked extended shifts. Interns on a traditional schedule with extended work shifts (24+ hours) made 35.9% more serious medical errors than on an intervention schedule that eliminated extended shifts (136.0 vs 100.1 per 1000 patient-days), including 56.6% more non-intercepted serious errors. On the critical care units, the total rate of serious errors (all staff) was 22.0% higher on the traditional schedule (193.2 vs 158.4 per 1000 patient-days), and interns' serious medication errors were 20.8% higher. Most strikingly, interns made 5.6 times as many serious diagnostic errors on the traditional schedule (18.6 vs 3.3 per 1000 patient-days). [2]

Complementing direct-observation studies, a large national prospective survey of 2,737 interns (17,003 monthly reports) quantified a steep dose-response between frequency of extended-duration shifts (≥24 h) and self-reported fatigue-attributed harm. Compared with months with no extended shifts, the odds of reporting at least one fatigue-related significant medical error were 3.5 (95% CI 3.3 to 3.7) with 1 to 4 extended shifts per month and 7.5 (95% CI 7.2 to 7.8) with ≥5 extended shifts per month. Fatigue-related preventable adverse events also rose sharply (OR 8.7 and 7.0, respectively), and reported fatality-associated errors were higher in the ≥5 extended-shifts group (OR 4.1, 95% CI 1.4 to 12). [3]

Fatigue risk is not confined to in-hospital decisions; it affects "after-action" safety relevant to staffing policies in regulated operations. A New England Journal of Medicine study on interns' commuting outcomes found the odds of a motor vehicle crash after an extended shift were more than double, and near-miss incidents were more than five times as likely after an extended shift versus a non-extended shift. [4]

In nursing, where regulated medication administration and documentation tasks share structural similarity with GxP batch record and release decisions, NIOSH summarizes evidence that nurses had over three times the odds of making an error when working 12 or more hours compared with 8.5-hour shifts, and that the risk for patient care errors almost doubled when critical care nursing shifts lasted longer than 12.5 hours. [5]

Time-on-task and time-awake effects provide a mechanistic bridge from "hours worked" to "error probability." Sleep loss reliably impairs vigilant attention, producing measurable "lapses" (failures to respond) that emerge after prolonged wakefulness and peak across circadian low points; one influential review illustrates this with a 42-hour total sleep deprivation protocol in which response failures began after about 16 hours awake and peaked after about 26 hours awake. In operational terms, this matters because many real-world oversight tasks are vigilance-heavy: they require the operator to detect rare but critical anomalies in a stream of "normal" system behavior. [6]

Evidence is also available in surgical performance, a close analog to "physical AI" settings because actions are often irreversible and occur under time pressure, multi-modal monitoring, and high consequence. A systematic review summarized that sleep deprivation negatively impacts technical performance in standardized simulated settings by about 11.9% to 32% on balance. [7] A broader systematic review of fatigue effects in surgeons reported heterogeneity but found that 35.4% of real-life studies observed deterioration in surgical outcomes when fatigue was present. [8]

Direct evidence specific to regulated laboratory technologist contexts is thinner, but available studies report that night shift fatigue and drowsiness were associated with lower alertness: a plausible precursor to increased analytical and documentation errors in regulated lab work, although error rate endpoints were not measured. [9]

Summary. The best-quantified clinical literature supports a defensible baseline claim: fatigue increases error rates and reduces attention, and it does so with large effect sizes (e.g., +36% serious errors; 3.5 to 7.5 times the odds of fatigue-related errors). [2] [3]

Cognitive load and AI oversight quality

Validated workload instruments matter because they allow "cognitive load" to be treated as a measurable hazard variable rather than an informal concept. The NASA Task Load Index (NASA-TLX) is a widely used subjective workload measure in human factors research. [10]

In healthcare decision support, multiple lines of evidence document automation bias (errors driven by over-reliance on automated recommendations) and quantify how wrong advice can degrade performance relative to unaided decisions. A systematic review of automation bias across domains (including a health care meta-analysis it summarizes) reports that incorrect decision support increased the risk of commission errors by 26% compared with when users did not have decision support, illustrating a measurable "harm of wrong help." [11] [16]

A particularly "GxP-shaped" clinical analog is electronic prescribing, where users approve or reject automated alerts. In a controlled experiment with 120 final-year medical students testing decision support in e-prescribing, incorrect decision support increased omission errors by 24.5 to 33.3 percentage points in some conditions, and incorrect alerts increased commission errors (accepting or acting on wrong advice) by 51.7 to 65.8 percentage points: a clear quantification of oversight failure when automation is wrong. Notably, the same study reported that task complexity and interruptions did not meaningfully change automation bias effects, implying that over-reliance can be resilient even without an explicit "overload" manipulation. [12]

More recent AI-era evidence in clinical expertise tasks shows that automation bias can persist even when AI improves mean performance. In a study of computational pathology judgments (tumor cell percentage estimation) with 28 trained pathology experts, adding AI increased overall performance but still produced a measurable automation bias rate of 7%, defined as initially correct evaluations overturned by erroneous AI advice. Time pressure did not increase the occurrence of automation bias, but appeared to increase its severity via stronger reliance on negative system consultations and consequent performance decline. [13]

Outside medicine, aviation and process-control studies provide quantified corroboration that workload, vigilance demands, and the difficulty of verification can degrade oversight. UK Civil Aviation Authority Paper 2004/10, citing earlier research, reports that pilots erred on 55% of occasions when the automation gave incorrect information, even though correct cross-check information was available: an outcome consistent with "verification failure" under operational complexity. [14]

Oversight failure is not only a function of overload. Neuroergonomics work maps undesirable neurocognitive states that come before performance loss, including mind wandering, effort withdrawal, perseveration and inattentional phenomena, and notes that passive monitoring can promote mind wandering, which is relevant to supervising highly automated lab robotics where long stretches are uneventful. [15]

Summary. The strongest quantified evidence does not support a simplistic "higher cognitive load, more automation bias" monotone curve in all settings. Instead, it supports three points: (1) wrong automation can measurably worsen decisions even when average performance improves, (2) supervision tasks are fragile under both overload and underload, and (3) verification complexity and poor human-machine interface design can dominate the failure mode. [12]

Automation bias and human-AI failure compounding

A central question is whether the combination "degraded human + AI" can produce outcomes worse than either alone. The strongest empirical support for compounding risk comes from automation bias evidence showing that wrong recommendations can flip correct human judgments into incorrect actions (negative consultations) and yield large commission error rates, and from fatigue literature showing large increases in attentional failure and error probability. [12]

In healthcare, a systematic review of automation bias documents measurable "negative consultation rates" (correct unaided answers changed to incorrect after consulting a decision support system) in the range of about 6 to 11% in some clinical question contexts, and summarizes that incorrect decision support increases commission error risk in meta-analytic estimates (risk ratio 1.26, 95% CI 1.11 to 1.44, from four studies). This establishes a quantifiable pathway for harm amplification: even if average accuracy improves, a nontrivial subset of cases become worse specifically because automation was consulted. [16]

The e-prescribing experiment described earlier quantifies this failure mode in especially stark terms: when decision support was incorrect, commission errors (acting on the incorrect cue) rose by 51.7 to 65.8 percentage points, and omission errors (failing to act when necessary) rose by 24.5 to 33.3 percentage points in measured conditions. This pattern is the "human-AI failure compounding" signature relevant to a physical AI lab: the AI cue drives the human to authorize an action that is wrong, and oversight fails at scale when the cue is wrong. [12]

Aviation literature provides complementary evidence that even trained professionals can fail to detect automation anomalies in the presence of cross-checkable data, with synthesized findings that errors occurred on about 55% of occasions when automation presented incorrect information. This supports the proposition that "just tell operators to verify" is not a sufficient control when verification is cognitively costly, time-pressured, or vigilance-degraded. [14]

How does degraded operator state interact with automation misuse? Evidence is mixed. One study on automated decision aids under sleep loss found decision aids could help maintain performance during sleep loss and reported that sleep-deprived participants were less prone to complacency and automation bias in that experimental setting, while their performance on a secondary task declined, suggesting that none of this can be assumed without domain- and interface-specific validation. [17]

However, even without a universal "fatigue increases automation bias" law, the compounding risk argument remains evidence-based because fatigue measurably increases attentional failures and error rates in high-stakes work, and automation bias measurably produces commission and negative-consultation errors when automation is wrong. The compound risk arises whenever (a) the AI can be wrong in rare but high-impact ways, and (b) the human's detection and verification capacity is impaired, regardless of whether the impairment increases "trust" or simply reduces detection bandwidth. [2]

Compounding risk in physical AI environments

The surgical fatigue literature shows that sleep deprivation measurably degrades technical skill in simulation (an estimated 11.9 to 32% decrement), and a sizeable fraction of real-life studies report worse outcomes under fatigue. This is not identical to lab automation approval workflows, but it is highly analogous in consequence structure: small degradations in precision, vigilance, or judgment can produce irreversible harm. [7]

A narrative review of partially automated driving reports slower responses when drivers supervise partial automation, and smaller vigilance decrements in manual driving. [18] By extension, and as this review's own inference rather than a finding of that study, the same pattern of long low-event intervals punctuated by rare critical decisions applies to approving AI-driven physical actions in automated labs.

Summary. The evidence supports treating operator state as a measurable hazard variable in physical automation settings and treating human oversight as a control that can fail systematically. Other safety-critical domains handle known human limitations in three ways: they reduce exposure (work-hour controls), add layers (two-person verification), or restrict action when physiological indicators suggest impairment (alcohol interlocks, drowsiness detection alarms). [19]

Regulatory and guidance framework

Cross-regulator AI governance expectations

In January 2026, FDA and EMA published "Guiding Principles of Good AI Practice in Drug Development," defining 10 principles intended to inform practice across the drug lifecycle (including manufacturing and post-marketing). Several principles are relevant to human oversight, especially the explicit expectation that performance assessment evaluates the complete system including human-AI interactions, and that a risk-based approach applies proportionate validation, mitigation and oversight based on context of use and model risk. The principles are high-level, as expected for cross-regulator alignment. [1]

EU AI Act requirements for high-risk systems

The EU AI Act (Regulation (EU) 2024/1689) is now the most explicit binding law grounding systemic risk management and human oversight expectations for high-risk AI. Article 9 requires a documented, maintained risk management system that is a "continuous iterative process planned and run throughout the entire lifecycle" and includes the identification and analysis of the known and the reasonably foreseeable risks to health, safety or fundamental rights. Article 14 requires that high-risk AI systems be designed and developed so they can be "effectively overseen" by natural persons during use, that oversight aims to prevent or minimise risks, and that oversight measures are commensurate with the risks, level of autonomy and context of use. [20]

Quality system and human factors anchors in U.S. regulation and standards

On the U.S. device side, FDA's Quality Management System Regulation (QMSR) modernizes 21 CFR Part 820 and explicitly incorporates ISO 13485:2016 by reference to harmonize quality system expectations. This anchors regulated expectations around documented processes, competence, corrective and preventive action, and risk-based management. [21]

FDA's guidance on applying human factors and usability engineering to medical devices recommends that device and system design reduce use-related risk through systematic usability engineering and validation. It provides a regulatory rationale for (a) treating operator limitations as foreseeable, and (b) implementing interface and workflow controls that reduce error when humans interact with devices. [22]

On the pharmaceutical quality system side, ICH Q10 defines a model for an effective pharmaceutical quality system across the product lifecycle. It reached Step 4 in June 2008, was adopted by EMA, and was issued by FDA as final guidance in April 2009. [23] (FDA) (EMA)

Physiological-state restrictions in safety-critical regulated practice

Multiple regulated sectors already restrict access or operational authority based on physiological state, most clearly where the marker is directly tied to impairment and has enforceable thresholds.

Aviation and transportation: alcohol thresholds

Aviation regulation explicitly prohibits acting as a crewmember while having an alcohol concentration of 0.04 or greater in blood or breath, and also prohibits acting within 8 hours of consuming alcohol. This is an unambiguous precedent: a regulated high-stakes domain uses a physiological marker with a numeric threshold to gate operational authority. [24]

On-road systems go further by implementing physical gating through alcohol ignition interlocks. CDC describes ignition interlocks as breath-test devices connected to vehicle ignition such that the vehicle will not start unless the driver's BAC is below a pre-set limit, "usually 0.02 g/dL." [25] The Community Preventive Services Task Force similarly describes interlocks as preventing operation when BAC is above a specified level, "usually 0.02% to 0.04%," making them one of the clearest real-world examples of physiological gating of an irreversible physical action (engine start). [29]

Nuclear power: fatigue management controls and self-declaration

In nuclear operations, the U.S. Nuclear Regulatory Commission mandates fatigue management elements within Fitness-for-Duty rules. NRC work-hour rules explicitly require scheduling "consistent with the objective of preventing impairment from fatigue" and impose controls such as limits on total hours (e.g., caps within a 7-day period) and required breaks, functioning as a structural gate that reduces exposure to fatigued work states. [19]

NRC rules also operationalize state-based gating behaviorally. For an individual working under a waiver of the work-hour limits who declares that they are not fit due to fatigue, the rules require a break of at least 10 hours before covered duties resume [26], and a fatigue assessment must be conducted in response to a self-declaration, except where the individual is given a 10-hour rest break. [27] This is not biometric, but it is state-based gating embedded in regulation: an operator state signal (self-report) is treated as sufficient to restrict duty in safety-critical work.

Fatigue monitoring systems with quantified outcome impact

Outside codified government rules, fatigue monitoring systems are deployed at scale in transport and heavy industry, and there is emerging randomized evidence that alarm and feedback change outcomes. In an Australian randomised cross-over trial of 24 fleet drivers, enabling Optalert's drowsiness alarms was associated with 34% fewer cautionary and 53% fewer critical drowsiness events than with alarms off (unadjusted rate ratios; the adjusted models were not statistically significant), and Seeing Machines alarms were associated with 55% fewer validated fatigue events, suggesting that monitoring plus intervention can alter hazardous event rates in real-world-like driving contexts. [28]

Sources

  1. FDA and EMA, Guiding Principles of Good AI Practice in Drug Development (January 2026). https://www.fda.gov/media/189581/download
  2. Effect of reducing interns' work hours on serious medical errors in intensive care units, New England Journal of Medicine 2004;351:1838-48. https://doi.org/10.1056/NEJMoa041406
  3. Impact of extended-duration shifts on medical errors, adverse events, and attentional failures, PLoS Medicine 2006 (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC1705824/
  4. Extended work shifts and the risk of motor vehicle crashes among interns, New England Journal of Medicine 2005;352:125-34. https://doi.org/10.1056/NEJMoa041401
  5. NIOSH, Work schedules and long work hours training for nurses, module 3, patient care errors. https://www.cdc.gov/niosh/work-hour-training-for-nurses/longhours/mod3/10.html
  6. Lim and Dinges (2008), Sleep deprivation and vigilant attention, Annals of the New York Academy of Sciences. https://doi.org/10.1196/annals.1417.002
  7. Systematic review of sleep deprivation and surgical technical performance, The Surgeon 2020 (PubMed). https://pubmed.ncbi.nlm.nih.gov/32057670/
  8. Systematic review of fatigue effects in surgeons, British Journal of Surgery (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC10771255/
  9. Impact of fatigue, sleep quality and drowsiness on medical laboratory technologists' alertness during night shift, Jurnal Ilmiah Teknik Industri 23(1):52-61, 2024. https://doi.org/10.23917/jiti.v23i1.3029
  10. NASA Task Load Index (NASA Technical Reports Server). https://ntrs.nasa.gov/citations/20000021488
  11. Automation bias and verification complexity: a systematic review, Journal of the American Medical Informatics Association 2017 (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC7651899/
  12. Controlled experiment on automation bias in electronic prescribing, BMC Medical Informatics and Decision Making 2017 (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC5356416/
  13. Study of automation bias and time pressure in computational pathology, Bildverarbeitung für die Medizin 2025, pp. 129-134. https://doi.org/10.1007/978-3-658-47422-5_27 (preprint: https://arxiv.org/abs/2411.00998)
  14. UK Civil Aviation Authority, CAA Paper 2004/10, Flight crew reliance on automation. https://www.caa.co.uk/publication/download/12922
  15. A neuroergonomics approach to mental workload, engagement and human performance, Frontiers in Neuroscience 2020. https://doi.org/10.3389/fnins.2020.00268
  16. Automation bias: a systematic review of frequency, effect mediators, and mitigators, Journal of the American Medical Informatics Association 2012 (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC3240751/
  17. Human performance consequences of automated decision aids in states of sleep loss, Human Factors 2011. https://doi.org/10.1177/0018720811418222
  18. Narrative review of supervision of partially automated driving, Frontiers in Psychology 2021 (PMC). https://pmc.ncbi.nlm.nih.gov/articles/PMC8081833/
  19. 10 CFR 26.205, Work hours (eCFR). https://www.ecfr.gov/current/title-10/chapter-I/part-26/subpart-I/section-26.205
  20. Regulation (EU) 2024/1689, the EU Artificial Intelligence Act (EUR-Lex). https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng
  21. FDA, Quality Management System Regulation (QMSR). https://www.fda.gov/medical-devices/postmarket-requirements-devices/quality-management-system-regulation-qmsr
  22. FDA guidance, Applying Human Factors and Usability Engineering to Medical Devices. https://www.fda.gov/regulatory-information/search-fda-guidance-documents/applying-human-factors-and-usability-engineering-medical-devices
  23. ICH Q10, Pharmaceutical Quality System. https://database.ich.org/sites/default/files/Q10%20Guideline.pdf ; FDA guidance page https://www.fda.gov/regulatory-information/search-fda-guidance-documents/q10-pharmaceutical-quality-system ; EMA page https://www.ema.europa.eu/en/ich-q10-pharmaceutical-quality-system-scientific-guideline
  24. 14 CFR 91.17, Alcohol or drugs (eCFR). https://www.ecfr.gov/current/title-14/chapter-I/subchapter-F/part-91/subpart-A/section-91.17
  25. CDC, Ignition interlocks. https://www.cdc.gov/impaired-driving/ignition-interlock/index.html
  26. 10 CFR 26.209, Self-declarations (eCFR). https://www.ecfr.gov/current/title-10/chapter-I/part-26/subpart-I/section-26.209
  27. 10 CFR 26.211, Fatigue assessments (eCFR). https://www.ecfr.gov/current/title-10/chapter-I/part-26/subpart-I/section-26.211
  28. Evaluation, validation and comparison of fatigued driving monitoring systems, final report, May 2024 (Australian Automobile Association). https://www.aaa.asn.au/wp-content/uploads/2024/10/Evaluation-Validation-Comparison-of-Fatigued-Driving-Monitoring-Systems.pdf
  29. Community Preventive Services Task Force, Motor vehicle injury: alcohol-impaired driving, ignition interlocks (The Community Guide). https://www.thecommunityguide.org/findings/motor-vehicle-injury-alcohol-impaired-driving-ignition-interlocks.html