Scope: target ID, hit discovery, lead optimization, de novo generation, protein/antibody design, foundation models, virtual screening, ADMET, medicinal chemistry, candidate prioritization, experiment planning, protocol generation, literature synthesis, closed-loop discovery.
Headline finding: AI in drug discovery has crossed the line from "recommends" to "acts." Documented systems now (a) autonomously plan and execute physical syntheses on robots (Coscientist, ChemCrow/RoboRXN), (b) run about 2.2 million automated imaging experiments per week whose results AI models analyze (Recursion OS), (c) run end-to-end agentic closed loops from target ID to wet-lab execution with only three human confirmation points (Insilico LabClaw/LifeStar2), and (d) have produced an AI-discovered and AI-designed drug with positive Phase IIa data that is now in Phase III (Insilico's rentosertib). Yet the regulatory record shows a governance gap precisely at the discovery/action boundary: FDA's Jan 2025 AI draft guidance explicitly excludes drug discovery from its credibility framework, and no predicate rule (GLP/GCP/GMP) governs AI-driven decisions until GLP tox studies begin. The accountable human at most discovery action points today is an individual scientist or an internal committee whose decision rationale is captured, at best, in an ELN entry or meeting minutes. Counter-evidence is real: human-in-the-loop checkpoints are standard in commercial systems, and several high-profile AI programs failed despite (or because of) the AI. Human review exists, but it is informal and subject to documented automation bias.
1. Workflow-by-Workflow: The Chain from AI Output to Real-World Action
1.1 Target identification (AI nominates a biological target → multi-year program + money + animal work committed)
Chain: AI target engine (PandaOmics, BenevolentAI Knowledge Graph, Recursion phenomics maps) ranks a novel target → program team accepts nomination → chemistry campaign staffed, CRO contracts signed, in vivo target-validation studies initiated. Consequence: money committed; animals used; the entire downstream program inherits the AI's choice. Accountable human today: VP/Head of Biology or a Target Review Committee. Evidence artifact: internal target nomination memo/presentation; occasionally a peer-reviewed paper years later. No regulator-required signed record at this point.
Evidence: - Claim: Insilico's PandaOmics identified TNIK as the top-ranked fibrosis target from multi-omics integration, and that AI-nominated target became the foundation of a program now in Phase III. / Source: Insilico Medicine press release (Phase III initiation) / URL: https://insilico.com/news/xmjsn4l091-insilico-initiates-phase-iii-clinical-tr / Date: 2026-07-07 / Excerpt: "In the Nature Biotechnology paper, TNIK was reported as the top-ranked candidate in the protein and receptor kinase discovery scenario... Rentosertib... was discovered and designed through Insilico's Pharma.AI platform. The program combines a novel fibrosis target prioritized by Biology42: PandaOmics..." / Context: Primary company disclosure; underlying paper is Ren et al., "A small-molecule TNIK inhibitor targets fibrosis in preclinical and clinical models," Nature Biotechnology (2024). / Confidence: High. - Claim: Recursion runs ~2.2 million automated phenotypic experiments per week; AI models then identify disease-relevant cellular signatures and drugs that reverse them. / Source: ClinicalMetric census + Recursion disclosures / URL: https://clinicalmetric.com/insights/ai-drug-discovery-human-trials-2026 / Date: 2026-06-09 / Excerpt: "Their operating system runs approximately 2.2 million experiments per week using automated imaging of cellular phenotypes under thousands of small molecule and genetic perturbations." / Context: Recursion's December 2021 collaboration and license agreement with Roche and Genentech carried a $150 million upfront payment (Recursion Form 8-K). / Confidence: Medium-high (company-reported figure, widely repeated). - Claim: AstraZeneca has deployed a production agentic system (ChatInvent) for real-world discovery workflows. / Source: He et al., "Democratising real-world drug discovery through agentic AI," Drug Discovery Today 31(2):104605 / URL: https://pubmed.ncbi.nlm.nih.gov/41548711/ (DOI https://doi.org/10.1016/j.drudis.2026.104605) / Date: Jan 2026 / Excerpt: "an agentic system called ChatInvent, which has been integrated into the discovery pipeline at AstraZeneca to aid in molecular design and synthesis planning" / Context: Peer-reviewed; the authors write that the literature lacked examples of real-world adoption of such systems in drug discovery. / Confidence: High that the paper exists and reports deployment; details require the full paper.
1.2 De novo molecule generation & virtual screening (AI designs molecules → synthesis/purchase order placed)
Chain: Generative model (Chemistry42, MolMIM, GEMS, Centaur Chemist) proposes structures → models rank/filter by predicted potency/ADMET/synthesizability → a human (or an agent) triggers a synthesis order or a compound purchase from a vendor library (e.g., Enamine). Consequence: money spent, physical material created, chemistry labor committed. Accountable human today: medicinal/computational chemist; the "make" decision is often one scientist's call in an ELN. Evidence artifact: ELN entry, registration in corporate compound database.
Evidence: - Claim: Insilico's end-to-end pipeline went from target nomination to first-in-human in ~30 months for ~$2.6M in discovery costs. / Source: intuitionlabs report citing Luo et al. 2024 + Insilico disclosures / URL: https://intuitionlabs.ai/articles/ai-target-discovery-cns-drug-development / Date: 2026-04-14 / Excerpt: "INS018-055 proceeded from target nomination to first-in-human trials in ≈30 months at only about $2.6 million cost." / Context: Discovery-to-clinic path published in Nature Biotechnology (2024). / Confidence: Medium (cost figure is company-reported; timeline corroborated). - Claim: Exscientia/Sumitomo's DSP-1181, the first AI-designed drug in human trials, required <12 months of exploratory research vs ~4.5 years typical. / Source: BioSpectrum Asia (announcement coverage) / URL: https://www.biospectrumasia.com/news/50/15362/sumitomo-dainippon-exscientia-begin-clinical-study-of-ai-based-ocd-drug.html / Date: 2020-02-03 / Excerpt: "This project was delivered by the strong synergy of the joint research, requiring less than 12 months to complete the exploratory research phase, just a fraction of the typical average of 4.5 years using conventional research techniques." / Context: Phase I initiated in Japan Jan 2020; later discontinued (see §4). / Confidence: High for the announcement; the timeline is company-reported. - Claim: Recursion's REC-1245 moved "from target identification to IND enabling studies in under 18 months" with ~200 compounds synthesized. / Source: intuitionlabs SDL report citing Recursion / URL: https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd / Date: 2026-08-05 / Excerpt: "the company reports the program 'moved from target ID to IND enabling studies in under 18 months with approximately 200 compounds synthesized'" / Context: FDA cleared the IND; REC-1245 (RBM39 degrader) in Phase I/II. / Confidence: Medium-high (company-reported, IND clearance is externally verifiable).
1.3 Autonomous synthesis & experiment execution (AI plans → robot physically executes)
This is the clearest documented case of AI decision → physical action without step-level human approval.
Evidence: - Claim: Coscientist, a GPT-4-driven system, autonomously designed, planned and executed real chemistry experiments including optimizing palladium-catalyzed cross-couplings on robotic liquid-handling hardware. / Source: Boiko, MacKnight, Kline, Gomes, Nature 624:570-578 / URL: https://www.nature.com/articles/s41586-023-06792-0 (DOI 10.1038/s41586-023-06792-0) / Date: 2023-12-20 / Excerpt: "we show the development and capabilities of Coscientist, an artificial intelligence system driven by GPT-4 that autonomously designs, plans and performs complex experiments by incorporating large language models empowered by tools such as internet and documentation search, code execution and experimental automation." / Context: Tier-1 peer-reviewed. The paper raises "potential dual-use consequences" and provides a brief dual-use study in its Supplementary Information ("Safety implications: Dual-use study"). / Confidence: High. - Claim: ChemCrow (GPT-4 + 18 chemistry tools) autonomously planned and executed syntheses of an insect repellent (DEET) and three organocatalysts on IBM's cloud-connected RoboRXN platform. / Source: Bran et al., Nature Machine Intelligence 6:525-535 / URL: https://www.nature.com/articles/s42256-024-00832-8 (preprint: https://arxiv.org/pdf/2304.05376v5) / Date: 2024-05-08 / Excerpt: "Using RoboRXN, ChemCrow autonomously ran the syntheses of an insect repellent (DEET) and three known thiourea organocatalysts" / Context: Notably, the paper also documents that predicted procedures "are not always directly executable on the RoboRXN platform," and that ChemCrow iteratively adapted the procedure "(such as increasing solvent quantity) until the synthesis procedure is fully valid, thereby removing the need for human intervention." / Confidence: High. - Claim: Insilico's LabClaw (announced May 2026) is a five-agent "Agent-Guard" system operating its fully automated LifeStar2 lab (six functional islands, dozens of devices, AGV material transfer), closing the loop from target discovery through wet-lab execution and report generation, with human confirmation only "at critical junctures." In the disclosed case study, researchers made decision confirmations at only three points in an end-to-end run. / Source: Insilico press release via EurekAlert / URL: https://www.eurekalert.org/news-releases/1127117 / Date: 2026-05-06 / Excerpt: "the system incorporates a Human-in-the-Loop approval mechanism at critical junctures... Throughout the entire process, researchers only need to make decision confirmations at three critical junctures, with all other repetitive, foundational experimental work automatically completed by the system" / Context: Primary vendor disclosure (self-reported; no independent audit of what the "three junctures" verify). The HITL checkpoints exist, but their content, authority basis and evidentiary record are not publicly defined. / Confidence: High that the system and claims exist; low on what the checkpoints actually enforce. - Claim: AstraZeneca's iLab (Gothenburg) is a prototype fully automated medicinal chemistry lab where "once the compounds have been tested, AI steps in to analyse the data and suggest new compounds to make and test." / Source: AstraZeneca R&D page / URL: https://www.astrazeneca.com/r-d/our-technologies/ilab.html / Date: accessed 2026 / Excerpt: "Once the compounds have been tested, AI steps in to analyse the data and suggest new compounds to make and test." / Context: Note the verb: AI "suggests"; AstraZeneca's public framing keeps humans selecting. Contrast with Insilico's "autonomous coordination" framing. / Confidence: High. - Counter-case: Eli Lilly's cloud lab (Lilly Life Sciences Studio, part of a ~$90M 2017 San Diego investment, operated by Strateos) was abandoned in 2024 and equipment sold. / Source: C&EN / URL: https://cen.acs.org/physical-chemistry/computational-chemistry/Self-driving-labs-changing-chemists/104/web/2026/06 / Date: 2026-06 / Excerpt: "The pharma giant gave up on that dream in 2024, when it disassociated itself from Strateos... 'it never really worked out, says Thomas Fessard, cofounder of SpiroChem.'" / Context: Important disconfirming evidence for autonomy hype. Separately, in January 2026 NVIDIA and Lilly announced they will invest up to $1 billion over five years in an AI co-innovation lab connecting Lilly's "agentic wet labs" with computational dry labs (NVIDIA, 12 Jan 2026). / Confidence: High.
1.4 ADMET prediction → in vivo / animal-study decisions
Chain: In silico ADMET/PK models (e.g., ADMET Predictor, Insilico ADMET.ai, Bayer's platform) triage thousands of compounds → only predicted-clean compounds proceed to animal PK/tox studies. Consequence: which animals are dosed, which compounds die; the FDA's Jan 2025 guidance explicitly contemplates AI reducing the number of nonclinical animal studies, i.e., AI output substituting for empirical safety data. Accountable human: DMPK/tox lead; IACUC approves animal protocols. Evidence artifact: study protocol + ELN.
Evidence: - Claim: FDA's 2025 draft guidance names "reducing animal-based pharmacokinetic and toxicologic studies" as an in-scope AI use requiring credibility assessment. / Source: Alethium analysis of FDA draft guidance / URL: https://www.alethium.health/resources/fdas-ai-credibility-framework-what-the-january-2025-draft-guidance-requires / Date: 2026-03-26 / Excerpt: "The guidance explicitly names these in-scope use cases: reducing animal-based pharmacokinetic and toxicologic studies, predictive modeling for clinical pharmacokinetics..." / Context: Corroborated by PointCross: "the guidance specifically enables AI to reduce the number of nonclinical pharmacokinetic, pharmacodynamic, and toxicologic studies required." / Confidence: High. - Claim: DSP-1181 was halted after Phase I (MDPI review); one 2026 review attributes the discontinuation to a QT-prolongation signal in animal studies. / Source: Frontiers in Pharmacology review / URL: https://www.frontiersin.org/journals/pharmacology/articles/10.3389/fphar.2026.1870527/full / Date: 2026-07-03 / Excerpt: "in 2022, the compound was discontinued during Phase I after preclinical safety studies revealed a prolonged QT interval in animal models... accelerated hit identification does not guarantee clinical progression and underscores the need for better integration of safety and pharmacokinetic prediction into AI-driven design pipelines." / Context: The cause rests on this one secondary review; other reviews say only that the compound was halted after Phase I. / Confidence: Medium (secondary review). - Claim: ADMET Predictor is used as a "'Tier Zero' screening tool... to triage thousands of compounds virtually before synthesis or in vivo testing." / Source: Pharmaron knowledge center (CRO poster summary) / URL: https://www.pharmaron.com/knowledge-center/admet-predictor-in-silico-ml/ / Date: 2026-02-24 / Context: Tier-3-ish but reflects standard industry practice; retrospective validation R²=0.9 dose prediction on 113 marketed drugs. / Confidence: Medium.
1.5 Protein/antibody design (generative models → physical constructs cloned, expressed, tested in animals)
Chain: RFdiffusion / ProteinMPNN / AbCellera / Xaira pipelines generate binder or antibody sequences de novo → in silico filters (AF2-multimer confidence) rank them → genes synthesized, proteins expressed, binding validated, top candidates go to animal studies. Accountable human: protein scientist/PI. Evidence artifact: sequence records, ELN; no regulated record until preclinical development.
Evidence:
- Claim: RFdiffusion fine-tuned for antibodies produced de novo antibodies binding user-specified epitopes "with atomic precision," confirmed by cryo-EM against influenza hemagglutinin, a C. difficile toxin, and PHOX2B. / Source: Genetic Engineering & Biotechnology News on Bennett et al., Nature (2025) / URL: https://www.genengnews.com/topics/artificial-intelligence/ai-designed-antibodies-achieve-atomic-precision-to-enhance-drug-discovery/ / Date: 2025-11-05 / Excerpt: "binding poses of antibody designs were confirmed by cryo-electron microscopy (cryo-EM) for an array of therapeutically relevant targets, including hemagglutinin... a potent toxin produced by the bacteria, Clostridium difficile, and PHOX2B" / Context: Co-authors founded Xaira Therapeutics ($1B committed capital, launched Apr 2024, board incl. former FDA head Scott Gottlieb). / Confidence: High.
- Claim: NVIDIA's BioNeMo Agent Toolkit packages biomolecular models (folding, docking, generative chemistry, protein design) as "callable skills" for autonomous AI agents, e.g. a generative_protein_binder_design meta-skill chaining RFdiffusion → ProteinMPNN → OpenFold3. / Source: MarkTechPost / URL: https://www.marktechpost.com/2026/06/29/nvidia-bionemo-agent-toolkit-turns-biomolecular-models-into-callable-skills-for-ai-agents-in-drug-discovery/ / Date: 2026-06-29 / Context: Infrastructure trend: the industry is deliberately converting models into agent-invokable tools, the exact substrate on which autonomous action chains are built. / Confidence: High.
1.6 Closed-loop discovery (the DMTA loop itself is being automated)
Evidence: - Claim: SDLs "close the design-make-test-analyze (DMTA) loop itself, using Bayesian optimization, active learning, reinforcement learning, or large language model (LLM) agents to choose the next experiment." / Source: intuitionlabs SDL report citing Nature Synthesis and a 2026 Nature Reviews Chemistry survey / URL: https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd / Date: 2026-08-05 / Excerpt: "algorithms propose, execute and interpret experiments with limited human intervention" / Context: Definitions from Tier-1 journals. / Confidence: High for definition. - Claim: The flagship academic SDL, Berkeley's A-Lab (originally reported as 41 novel compounds from 58 targets in 17 days; the corrected article now reports 36 compounds from 57 targets), received a formal Nature author correction (19 Jan 2026) after outside chemists (Palgrave et al.) challenged the novelty claims. The authors acknowledged that "the original claims of material novelty were subject to misinterpretation." / Source: Nature 624:86-91 and its author correction, plus jimmyresearch summary / URL: https://www.nature.com/articles/s41586-023-06734-w ; correction: https://www.nature.com/articles/s41586-025-09992-y ; https://jimmyresearch.com/modules/aifs-chemistry/ / Date: 2023-11-29 (article); 2026-01-19 (correction) / Excerpt: "Nature issued a correction, but critics maintain... that core concerns about whether genuinely new materials were made remain unresolved." / Context: Critical cautionary evidence: the novelty claims of an autonomous lab's flagship paper had to be corrected, and the problems were raised by outside chemists after publication. / Confidence: High. - Claim: A peer-reviewed regulatory analysis identifies a structural conflict between continuously learning SDL models and GMP's fixed, validated procedures: "continuous learning and opacity of decision-making processes" vs "the principles of GMP." / Source: Niazi, "Regulatory Perspectives for AI/ML Implementation in Pharmaceutical GMP Environments," Pharmaceuticals 18(6):901 (2025), cited by intuitionlabs / URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12195787/ and https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd / Date: 2025-06-16 / Excerpt: "The unique characteristics of AI/ML, such as continuous learning and opacity of decision-making processes, pose significant challenges to the principles of GMP" / Context: Same tension applies to discovery-phase GLP transition. / Confidence: Medium-high.
1.7 Literature synthesis / experiment planning / protocol generation (agentic co-scientists)
Evidence: - Claim: A research ecosystem of discovery agents now exists: PharmAgents (a virtual pharma built from multiple LLM agents), Google's TxGemma (therapeutics-specialized models), DiscoVerse ("multi-agent pharmaceutical co-scientist for traceable drug discovery"), CRISPR-GPT (agentic automation of gene-editing experiments, Nature Biomedical Engineering 2025). / Source: arXiv survey (Mozi) reference list / URL: https://arxiv.org/html/2603.03655v1 / Date: 2025-2026 / Confidence: High that systems exist; medium on deployment depth.
2. Regulatory & Evidentiary Frame (what record is required, where the gaps are)
- Claim: FDA's first AI draft guidance (Jan 6-7, 2025; Docket FDA-2024-D-4689) establishes a 7-step credibility assessment (question of interest → context of use → risk = model influence × decision consequence → plan → execute → document → adequacy/lifecycle), informed by >500 AI-containing submissions (2016-2023) and >800 comments. / Source: FDA CDER AI page; Federal Register notice of availability (7 Jan 2025) / URL: https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development and https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug ; seven-step framework: https://www.alethium.health/resources/fdas-ai-credibility-framework-what-the-january-2025-draft-guidance-requires / Date: accessed 2026-05 / Excerpt: "The content of this draft guidance was informed by... (3) CDER's experience with over 500 submissions with AI components from 2016 to 2023" / Confidence: High.
- Claim: The framework "explicitly excludes drug discovery and internal operational efficiencies from its reach." / Source: Drug Discovery News / URL: https://www.drugdiscoverynews.com/regulating-ai-in-drug-discovery-what-fda-ema-and-ich-guidance-means-for-pharma-r-d-17366 / Date: 2026-07-21 / Excerpt: "the framework explicitly covers the nonclinical, clinical, postmarketing, and manufacturing phases, and just as explicitly excludes drug discovery and internal operational efficiencies from its reach, at least for now." / Context: Corroborated by PointCross: "Drug discovery applications and purely operational AI tools are excluded from scope." / Confidence: High.
- Claim: EMA's Reflection Paper on AI in the Medicinal Product Lifecycle (adopted by CHMP 2024-09-09) covers AI "at any step of a medicines' lifecycle, from drug discovery to the post-authorisation setting," states "a human-centric approach should guide all development and deployment of AI and ML," and, for high-risk unqualified models in trials, expects "the full model architecture, logs from model development, validation and testing, training data and description of the data processing pipeline" in the dossier. / Source: intuitionlabs clinical-trials census quoting EMA; EMA reflection paper / URL: https://intuitionlabs.ai/articles/ai-clinical-trials-clinicaltrials-gov-census and https://www.ema.europa.eu/en/documents/scientific-guideline/reflection-paper-use-artificial-intelligence-ai-medicinal-product-lifecycle_en.pdf / Date: 2026-08-06 / Context: EMA's scope includes discovery, unlike FDA's; but it is a reflection paper (non-binding) and its discovery-phase expectations are principles, not inspection-enforced records. Joint FDA/EMA "Guiding Principles of Good AI Practice in Drug Development" (Jan 2026, 10 principles incl. "Human-centric by design") is likewise non-binding. / Confidence: High.
- Claim: 21 CFR Part 11 §11.10(e) requires secure, computer-generated, time-stamped audit trails that independently record the date and time of operator entries and actions that create, modify or delete electronic records, without obscuring previously recorded information; §11.50 requires signature manifestation, and §11.70 requires signatures to be linked to their records so that they "cannot be excised, copied, or otherwise transferred to falsify an electronic record by ordinary means" (eCFR, 21 CFR Part 11). But Part 11 only applies when a predicate rule (GLP Part 58, GCP, GMP Part 211) requires the record. Early discovery has no predicate rule, so no Part 11 obligation attaches to AI-driven synthesis decisions, compound selections, or experiment queuing. / Source: Klyverity Part 11 analysis + govalidation FAQ / URL: https://klyverity.com/blog/part-11-audit-trail-requirements and https://govalidation.com/knowledge/21-cfr-part-11-faq/ / Date: 2026 / Excerpt: "Part 11 governs how those records may be maintained electronically; it does not define what records must exist. Whether a record must be kept is determined by the predicate rule" (govalidation) / Context: Synthesized finding: signed, audit-trailed records of decisions exist as a legal requirement only from GLP nonclinical studies onward; everything upstream, where AI autonomy is advancing fastest, runs on informal artifacts. / Confidence: High (regulatory-logic synthesis from well-established rules).
- Claim: FDA enforcement practice confirms the evidentiary emphasis: according to Klyverity, "When an FDA investigator sits down to review your electronic signature system," the audit trail is what they ask for first, and they check "can they reconstruct exactly what happened, who did it, and when, without taking anyone's word for it?" / Source: Klyverity / URL: https://klyverity.com/blog/part-11-audit-trail-requirements / Confidence: Medium (consultancy framing, but consistent with known inspection practice).
3. The Accountable Human & Evidence Artifact Today (by transition point)
| Transition (AI decision → action) | Accountable human today | Evidence artifact today | Mandated? |
|---|---|---|---|
| AI-nominated target accepted | VP Biology / target review committee | Nomination slide deck/memo; sometimes later paper | No |
| AI-generated compound selected for synthesis | Medicinal chemist | ELN entry; compound registry | No (non-GLP) |
| AI-planned route executed on robot | Lab scientist / automation engineer | Robot run log (vendor-specific) | No |
| Virtual screen ranking → purchase/synthesis order | Computational chemist | Order record | No |
| ADMET triage → compound killed or advanced to in vivo | DMPK lead; IACUC (for animal use) | Protocol | Partial (IACUC) |
| SDL agent queues next experiment (closed loop) | Nominal HITL approver at "critical junctures" | LIMS records; checkpoint definition undisclosed | No |
| Lead series → candidate nomination | Candidate selection committee / CSO | Committee minutes, nomination package | Company-internal only |
| Candidate → GLP tox / IND-enabling | Head of preclinical dev; QA (GLP) | GLP study records, Part 11 audit trails | Yes (21 CFR 58, Part 11) |
| IND → first-in-human | CMO; sponsor; FDA 30-day review | IND dossier | Yes |
| AI model supporting a submission | Sponsor | Credibility assessment plan/report (FDA 2025 draft) | Draft guidance only |
Key observation: every formally mandated artifact sits downstream of the point where AI autonomy is expanding fastest. The discovery-phase action points rely on ELN discipline and committee culture, not on mandated records.
4. Counter-Arguments and Disconfirming Evidence
- Human-in-the-loop is the commercial norm, not the exception. Insilico's LabClaw has explicit HITL confirmation "at critical junctures"; AstraZeneca's iLab frames AI as suggesting compounds; ChemCrow/Coscientist ran under researcher-defined objectives. [894][585]
- Existing regulated-phase infrastructure (Part 11, GLP audit trails, ELN/LIMS) already covers the consequential, patient-facing decisions. Discovery decisions are reversible: a bad synthesis wastes reagents, not patients. [23]
- AI drug discovery has high-profile failures, suggesting the autonomy threat is overblown: DSP-1181 was halted after Phase I; BenevolentAI's BEN-2293 missed its Phase IIa efficacy endpoints (BenevolentAI announcement, 5 April 2023), BenevolentAI cut 30% of its workforce after its lead drug did not work, and the company later delisted from Euronext Amsterdam (BenevolentAI, Mar 2025); Recursion reported limited efficacy for REC-994, contributing to pipeline restructuring; Recursion acquired Exscientia in 2024. If AI decisions were dangerously unsupervised, we'd see safety disasters, not just efficacy failures. / Source: MDPI Pharmaceuticals 19(6):916 / URL: https://www.mdpi.com/1424-8247/19/6/916 (full text also at https://pmc.ncbi.nlm.nih.gov/articles/PMC13304925/) / Date: 2026-06-10 / Confidence: High for the failure facts. [259]
- Automation bias research shows human checkpoints degrade in practice. A 2025 systematic review of 35 peer-reviewed studies describes automation bias as "the tendency to over-rely on automated recommendations" and finds that explanations "are often insufficient to improve decision accuracy or mitigate" it. A nominal HITL checkpoint is not equivalent to effective human authority. / Source: Romeo and Conti, "Exploring automation bias in human-AI collaboration," Springer AI & Society / URL: https://link.springer.com/article/10.1007/s00146-025-02422-7 / Date: 2025-07-03 / Confidence: High (for the general phenomenon; not drug-discovery-specific).
- Lilly's remotely operated, automated Life Sciences Studio lab, run with Strateos and part of a ~$90M 2017 San Diego investment, was given up in 2024, evidence that full autonomy may arrive more slowly than vendors claim. [255]
5. Dual-Use & Safety Dimension (severity anchor for the action-point map)
- Claim: Urbina et al. showed a generative model (MegaSyn), repurposed for toxicity instead of avoidance, produced ~40,000 predicted-toxic molecules including nerve-agent-like structures in under six hours. / Source: Urbina, Lentzos, Invernizzi, Ekins, "Dual use of artificial-intelligence-powered drug discovery," Nature Machine Intelligence 4:189-191 (2022) / URL: https://www.nature.com/articles/s42256-022-00465-9 (full text: https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/) / Excerpt: "In less than 6 hours after starting on our in-house server, our model generated forty thousand molecules that scored within our desired threshold. In the process, the AI designed not only VX, but many other known chemical warfare agents" / Context: When generative models are wired to robotic synthesis (ChemCrow/RoboRXN precedent), the authority question "who may cause this compound to be made?" becomes a hard safety boundary. Coscientist's authors included a brief dual-use study. / Confidence: High.
6. Action-Point Map: Top 10 AI-Decision → Real-World-Action Points in Drug Discovery (ranked by consequence severity)
- IND submission / first-in-human dosing decision for an AI-discovered candidate. Consequence: humans dosed; accountable: CMO/sponsor; artifact: IND dossier (regulated). Rentosertib is the live case (now Phase III). [244]
- Candidate nomination → GLP tox program initiation. Consequence: first regulated studies; large animal use; multi-$M commitment; accountable: dev-candidate committee/CSO; artifact: nomination package + GLP records (regulated from this point). [23]
- Autonomous robotic synthesis of AI-designed compounds (agent plans route, robot executes). Consequence: physical creation incl. potential hazardous/dual-use chemistry; accountable: automation scientist; artifact: run log only. Coscientist, ChemCrow/RoboRXN, Insilico LifeStar2/LabClaw. [444][894]
- Closed-loop experiment selection in SDLs (AI queues the next physical experiment incl. CRISPR editing, cell work). Consequence: continuous unattended physical action; accountable: nominal HITL approver at a few checkpoints (three in Insilico's disclosed LabClaw run); artifact: LIMS records, undisclosed checkpoint semantics. [894][17]
- AI/ADMET triage substituting for empirical tox studies (negative decisions). Consequence: fewer nonclinical animal studies run (FDA's draft guidance names "reducing animal-based pharmacokinetic and toxicologic studies" as an in-scope AI use); accountable: DMPK/tox lead; emerging artifact: credibility assessment report (draft guidance). [679]
- Target nomination. Consequence: a program and its funding committed to an AI-chosen biology; accountable: VP Biology/committee; artifact: memo/slides (unregulated). [244]
- Virtual-screen ranking → compound library purchase/synthesis orders. Consequence: compounds made and bought at scale (Recursion and Enamine curated screening libraries from "over 15,000 newly synthesized compounds"); accountable: computational chemist; artifact: order records. [17]
- De novo protein/antibody design → gene synthesis & animal testing. Consequence: novel bioactive proteins physically made (incl. toxin-binding constructs); accountable: protein scientist; artifact: sequence/ELN records. [591]
- Agentic experiment planning / protocol generation executed in cloud labs. Consequence: protocols run on shared robotic infrastructure (for example, Coscientist executed experiments through the Emerald Cloud Lab); accountable: remote researcher; artifact: platform run logs. [442]
- Literature-synthesis/hypothesis agents influencing program prioritization. Consequence: strategic direction shifted by agent-generated syntheses (hallucination risk propagates downstream); accountable: program lead.
Source List
- [2] FDA CDER, "Artificial Intelligence for Drug Development", https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development (Tier-1)
- [17] IntuitionLabs, "Self-Driving Labs in Pharma: Closed-Loop R&D That Works", https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd (2026-08; Tier-2 aggregator citing Nature Synthesis, Nature Reviews Chemistry, McKinsey, Grand View)
- [23] Klyverity, "21 CFR Part 11 Audit Trail Requirements", https://klyverity.com/blog/part-11-audit-trail-requirements (Tier-3 consultancy, consistent with regulation text)
- [35] Valkit / [453] GoValidation, Part 11 predicate-rule FAQ, https://govalidation.com/knowledge/21-cfr-part-11-faq/
- [54] ClinicalMetric, "AI-Discovered Drugs in Human Trials 2026", https://clinicalmetric.com/insights/ai-drug-discovery-human-trials-2026
- [159] Biosecurity Handbook, "AI-Enabled Pathogen Design", https://biosecurityhandbook.com/ai-biosecurity/ai-pathogen-design.html
- [240] IntuitionLabs, "How Many Clinical Trials Use AI? 2026 Census" (EMA reflection paper quotes), https://intuitionlabs.ai/articles/ai-clinical-trials-clinicaltrials-gov-census
- [244] Insilico Medicine, "Insilico Initiates Phase III Clinical Trial for Rentosertib", https://insilico.com/news/xmjsn4l091-insilico-initiates-phase-iii-clinical-tr (2026-07-07; primary)
- [245] Frontiers in Pharmacology, "AI in drug discovery: from algorithmic foundations to clinical translation", https://www.frontiersin.org/journals/pharmacology/articles/10.3389/fphar.2026.1870527/full (2026-07)
- [249] BioSpectrum Asia, "Sumitomo Dainippon, Exscientia begin clinical study of AI based OCD drug", https://www.biospectrumasia.com/news/50/15362/sumitomo-dainippon-exscientia-begin-clinical-study-of-ai-based-ocd-drug.html (2020-02-03)
- [255] C&EN, "Self-driving labs are changing how chemists work", https://cen.acs.org/physical-chemistry/computational-chemistry/Self-driving-labs-changing-chemists/104/web/2026/06 (Tier-2)
- [259] MDPI Pharmaceuticals 19(6):916, "AI in Drug Discovery: Clinical Failures, Regulatory Reality, and the Validation Crisis Behind the Hype", https://www.mdpi.com/1424-8247/19/6/916 (full text also at https://pmc.ncbi.nlm.nih.gov/articles/PMC13304925/) (2026-06; peer-reviewed)
- [442] Boiko, MacKnight, Kline, Gomes, "Autonomous chemical research with large language models" (Coscientist), Nature 624:570-578 (2023), https://www.nature.com/articles/s41586-023-06792-0 (Tier-1)
- [443][444] Bran et al., ChemCrow, arXiv 2304.05376 / Nature Machine Intelligence 6:525-535, https://www.nature.com/articles/s42256-024-00832-8 (preprint https://arxiv.org/pdf/2304.05376v5) (Tier-1)
- [445] BenevolentAI, top-line Phase IIa results for BEN-2293, 5 April 2023 (safety endpoint met; secondary efficacy endpoints not achieved), https://www.pharmiweb.com/press-release/2023-04-05/benevolentai-announces-top-line-phase-iia-results-for-its-topical-pan-trk-inhibitor-ben-2293-1-in-mild-to-moderate-atopic-dermatitis
- [450] Jimmy Research, "AI for Chemistry" (A-Lab controversy summary with Nature refs), https://jimmyresearch.com/modules/aifs-chemistry/
- [585] AstraZeneca, "The AstraZeneca iLab", https://www.astrazeneca.com/r-d/our-technologies/ilab.html (primary)
- [586] AICerts / [588] MarkTechPost, NVIDIA BioNeMo Agent Toolkit, https://www.marktechpost.com/2026/06/29/nvidia-bionemo-agent-toolkit-turns-biomolecular-models-into-callable-skills-for-ai-agents-in-drug-discovery/ (2026-06-29)
- [591] GEN News, "AI-Designed Antibodies Achieve Atomic Precision" (RFdiffusion antibodies, Nature 2025; Xaira), https://www.genengnews.com/topics/artificial-intelligence/ai-designed-antibodies-achieve-atomic-precision-to-enhance-drug-discovery/ (2025-11-05)
- [592] Urbina, Lentzos, Invernizzi, Ekins, "Dual use of artificial-intelligence-powered drug discovery," Nat. Mach. Intell. 4:189-191 (2022), https://www.nature.com/articles/s42256-022-00465-9 (full text https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/)
- [679] Alethium, "FDA's AI Credibility Framework: What the January 2025 Draft Guidance Requires", https://www.alethium.health/resources/fdas-ai-credibility-framework-what-the-january-2025-draft-guidance-requires
- [760] Drug Discovery News, "Regulating AI in drug discovery", https://www.drugdiscoverynews.com/regulating-ai-in-drug-discovery-what-fda-ema-and-ich-guidance-means-for-pharma-r-d-17366 (Tier-2)
- [763] PointCross, "AI in Nonclinical Reporting", https://pointcrosslifesciences.com/ai-in-nonclinical-reporting-from-hype-to-real-workflow/
- [771] Pharmaron, "ADMET Predictor: In Silico Screening", https://www.pharmaron.com/knowledge-center/admet-predictor-in-silico-ml/
- [894][896] Insilico Medicine / EurekAlert, "LabClaw the Intelligent System", https://www.eurekalert.org/news-releases/1127117 (2026-05-06; primary)
- [898] "Mozi: Governed Autonomy for Drug Discovery LLM Agents," arXiv 2603.03655, https://arxiv.org/html/2603.03655v1
- [906] Romeo and Conti, Springer AI & Society, "Exploring automation bias in human-AI collaboration: a review and implications for explainable AI" (2025), https://link.springer.com/article/10.1007/s00146-025-02422-7 (peer-reviewed)
- He et al., "Democratising real-world drug discovery through agentic AI" (ChatInvent), Drug Discovery Today 31(2):104605 (2026), https://pubmed.ncbi.nlm.nih.gov/41548711/
- Szymanski et al., "An autonomous laboratory for the accelerated synthesis of inorganic materials" (A-Lab), Nature 624:86-91 (2023), https://www.nature.com/articles/s41586-023-06734-w ; Author Correction (19 Jan 2026), https://www.nature.com/articles/s41586-025-09992-y
- Niazi, "Regulatory Perspectives for AI/ML Implementation in Pharmaceutical GMP Environments," Pharmaceuticals 18(6):901 (2025), https://pmc.ncbi.nlm.nih.gov/articles/PMC12195787/
- FDA, draft guidance "Considerations for the Use of Artificial Intelligence To Support Regulatory Decision-Making for Drug and Biological Products," Federal Register, 7 Jan 2025 (Docket FDA-2024-D-4689), https://www.federalregister.gov/documents/2025/01/07/2024-31542/considerations-for-the-use-of-artificial-intelligence-to-support-regulatory-decision-making-for-drug
- FDA and EMA, "Guiding Principles of Good AI Practice in Drug Development" (Jan 2026), https://www.fda.gov/media/189581/download
- EMA, "Reflection paper on the use of Artificial Intelligence (AI) in the medicinal product lifecycle," https://www.ema.europa.eu/en/documents/scientific-guideline/reflection-paper-use-artificial-intelligence-ai-medicinal-product-lifecycle_en.pdf
- 21 CFR Part 11 (eCFR), https://www.ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- Recursion Pharmaceuticals, Form 8-K (Roche/Genentech Collaboration and License Agreement, Dec 2021), https://www.sec.gov/Archives/edgar/data/1601830/000119312521349657/d242203d8k.htm
- NVIDIA, "NVIDIA and Lilly Announce Co-Innovation Lab to Reinvent Drug Discovery in the Age of AI" (12 Jan 2026), https://nvidianews.nvidia.com/news/nvidia-and-lilly-announce-co-innovation-lab-to-reinvent-drug-discovery-in-the-age-of-ai
- BenevolentAI, "EGM Results announcement" (delisting from Euronext Amsterdam, Mar 2025), https://www.benevolent.com/news-and-media/press-releases-and-in-media/egm-results-announcement/
