Skip to content
Regulayer™Human Control for AI
Book a live demo

Research library · Life sciences and GxP

The "set the menu" problem in drug discovery

Published evidence that in AI-led drug discovery the objective function, acquisition policy, sampling settings and training data determine which candidates are ever generated or tested. Downstream validation cannot reveal the alternatives that were never offered.

Compiled from public sources, 17 August 2026. Information, not legal advice.

An AI system increasingly determines which candidates are generated, surfaced or experimentally tested at all. If its objective or operating constraints change, downstream validation may never reveal the omitted alternatives, because the upstream system already determined the search space.

This report collects published, dated evidence on that question from peer-reviewed literature (the Nature family, J. Chem. Inf. Model., J. Cheminformatics, NeurIPS, J. Biomed. Inform.), official regulatory documents, and named documented cases. Excerpts are verbatim from the sources.

Two phenomena are separated throughout:

  • A, scientific and model bias: search-space determination that is scientifically suboptimal.
  • B, authority drift: an AI system acting under authority or constraints that no longer represent the current human decision.

Part I. Scientific and model bias: the model sets the menu suboptimally

A1. The MegaSyn toxin study: inverting one scoring sign re-determines the search space

Finding: Inverting a single scoring sign transformed a drug-discovery generator from a tool that avoids toxic chemical space into one that concentrates on it exclusively, generating 40,000 predicted-toxic molecules in under 6 hours, including VX and other known chemical warfare agents that were not in the training data.

Source: Urbina F., Lentzos F., Invernizzi C., Ekins S., "Dual use of artificial intelligence-powered drug discovery," Nature Machine Intelligence (2022). Published March 2022, peer-reviewed.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/

Excerpt:

"We simply proposed to invert this logic using the same approach to design molecules de novo, but now guiding the model to reward both toxicity and bioactivity instead... In less than 6 hours after starting on our in-house server, our model generated forty thousand molecules that scored within our desired threshold. In the process, the AI designed not only VX, but many other known chemical warfare agents that we identified through visual confirmation with structures in public chemistry databases... These new molecules were predicted to be more toxic based on the predicted LD50 in comparison to publicly known chemical warfare agents. This was unexpected as the datasets we used for training the AI did not include these nerve agents."

"Importantly, we had a human-in-the-loop with a firm moral and ethical 'don't-go-there' voice to intervene. But what if the human was removed or replaced with a bad actor? With current breakthroughs and research into autonomous synthesis, a complete design-make-test cycle applicable to making not only drugs, but toxins, is within reach."

Context: This is the clearest documented demonstration that the objective function, not the model architecture or the data, determines which region of chemical space is explored at all. The identical model, data and pipeline explored a completely disjoint region of molecular property space: the authors show a t-SNE plot in which the toxic molecules occupy a region entirely separate from the training distribution. Downstream validation on the original objective would never have surfaced these alternatives, and the reverse holds too. The second excerpt is also the strongest published articulation of phenomenon B: the authors frame the safeguard as a human authority check outside the model, and ask what happens when that human authority is removed.

Confidence: Very high. Peer-reviewed, Tier-1 journal, primary source.

A2. Reward hacking in molecular generative design: optimizers exploit the proxy, not the property

Finding: Generative molecular design guided by predictive models is "often susceptible to optimization failure due to reward hacking," where the optimizer produces molecules that score well because they fall outside the predictor's training distribution. The proxy diverges from the true property and the optimizer follows the proxy.

Source: Yoshizawa T., Ishida S., Sato T., Ohta M., Honma T., Terayama K., "A data-driven generative strategy to avoid reward hacking in multi-objective molecular design," Nature Communications 16, 2409 (11 March 2025). Peer-reviewed.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC11897179/ ; https://pubmed.ncbi.nlm.nih.gov/40069140/

Excerpt:

"However, this approach is often susceptible to optimization failure due to reward hacking, where prediction models fail to extrapolate, i.e., fail to accurately predict properties for designed molecules that considerably deviate from the training data."

"This phenomenon, well-known in reinforcement learning and game AI, occurs when optimization deviates unexpectedly from the intended goals... Consequently, the optimization process deviates from its intended direction. In such cases, the seemingly favorable predicted value of the designed molecule may be inaccurate. In fact, in drug design, there have been cases where unstable or complex molecules, distinct from existing drugs, have been designed due to reward hacking."

Context: Documents optimization converging on the wrong proxy in drug discovery. The paper's proposed fix, applicability-domain constraints, is a model-internal scientific mitigation: it lives inside the modeling layer.

Confidence: High. Tier-1 peer-reviewed.

A3. Failure modes in molecule generation: scoring-function exploitation that current metrics miss

Finding: Goal-directed molecular generation exploits bias in the scoring function that guides it, and these failure modes evade detection by the performance metrics in standard use.

Source: Renz P., Van Rompaey D., Wegner J.K., Hochreiter S., Klambauer G., "On failure modes in molecule generation and optimization," Drug Discovery Today: Technologies 32-33, 55-63 (December 2019). DOI: 10.1016/j.ddtec.2020.09.003. Restated in Langevin M., Vuilleumier R., Bianciotto M., "Explaining and avoiding failure modes in goal-directed generation of small molecules," Journal of Cheminformatics (1 April 2022), peer-reviewed.

URL: https://doi.org/10.1016/j.ddtec.2020.09.003 ; https://pmc.ncbi.nlm.nih.gov/articles/PMC8973583/

Excerpt, Renz et al., abstract:

"The evaluation of generative models remains challenging and suggested performance metrics or scoring functions often do not cover all relevant aspects of drug design projects. In this work, we highlight some unintended failure modes in molecular generation and optimization and how these evade detection by current performance metrics."

Excerpt, Langevin et al. on the Renz et al. result:

"The observed difference between Sopt, Smc and Sdc is explained in [12] by the fact that goal-directed generation algorithms exploit bias in the scoring function, defined as the presence of features that yield a high Sopt but do not generalize to Smc and Sdc."

Context: Search-space drift here is data-specific and model-specific: the optimizer colonizes whatever region the flawed surrogate rewards, which need not overlap with chemically or biologically meaningful space. The follow-up work finds that generated molecules exploit bias unique to the model they were optimized on, so the menu is partly an artifact of surrogate idiosyncrasy.

Confidence: High. The primary paper is peer-reviewed and widely cited, and the finding is restated in a later peer-reviewed paper that reproduced the experiment. The detailed excerpts quoted in an earlier draft came from a third-party blog summary and could not be found in the primary, so they were removed.

A4. Mode collapse and the "filtering trap": sampling settings silently narrow the design space

Finding: Common molecule-sampling settings (low top-k or top-p) cause mode collapse, that is valid but repetitive, low-diversity designs, consistently across architectures. Highly likely designs favour exploitation of known actives and sacrifice novelty, "limiting chemical space exploration."

Source: Özçelik R., Grisoni F., "How evaluation choices distort the outcome of generative drug discovery," Journal of Cheminformatics 17, 169 (12 November 2025), peer-reviewed. The preprint is arXiv:2501.05457, whose first version, submitted 24 December 2024, was titled "The Jungle of Generative Drug Discovery: Traps, Treasures, and Ways Out" and was retitled to match the journal version on 13 November 2025.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12613558/ ; https://arxiv.org/html/2501.05457v1

Excerpt:

"Considering a small token subset during molecule sampling (via top-k or top-p sampling) cause mode collapse (repetitive, low-diversity designs)."

"For top-k sampling, the smallest value of k (k=3) causes mode collapse, i.e., designs are valid but repetitive... We call this behavior, which is consistent across model architectures and training regimes, 'filtering trap'."

"Excessive likelihood. Highly-likely designs favor exploitation (similarity to known actives) but sacrifice novelty and diversity, limiting chemical space exploration."

Context: Search-space narrowing can be introduced by mundane configuration choices, a sampling constant, not only by malice or model failure. No downstream assay can reveal that an alternative candidate class was never generated because k was set to 3. A configuration parameter, changeable by an operator or by a software update, silently redefines the menu.

Confidence: High. Peer-reviewed, systematic across architectures.

A5. GuacaMol benchmark: GAN mode collapse documented, and trivially saturated benchmarks mislead

Finding: The canonical de novo design benchmark documented that ORGAN, a GAN, generated more than 50% invalid molecules with poor KL-divergence and zero FCD scores, which "might indicate mode collapse", and that several goal-directed tasks are saturated by simply picking molecules from ChEMBL, so "such trivial benchmarks are not suitable for the assessment of generative models."

Source: Brown N., Fiscato M., Segler M.H.S., Vaucher A.C., "GuacaMol: Benchmarking Models for De Novo Molecular Design," J. Chem. Inf. Model. (2019); arXiv:1811.09621. Preprint November 2018; journal 2019.

URL: https://arxiv.org/pdf/1811.09621

Excerpt:

"More than half the molecules it generates are invalid, and they do not resemble the training set, as shown by the KL divergence and FCD scores. This might indicate mode collapse, which is a phenomenon often observed when training GANs."

"From these results, it is evident that such trivial benchmarks are not suitable for the assessment of generative models for de novo molecular design. The ChEMBL database alone achieves excellent scores already and therefore these benchmarks cannot demonstrate the advantage of generative models over pure virtual screening."

Context: Baseline peer-reviewed evidence that mode collapse in chemical generative models is real and documented, and that the evaluation layer itself can fail to reveal that the generative system added nothing beyond the training distribution. That is a second sense in which downstream validation does not expose upstream menu-setting.

Confidence: High. Peer-reviewed, field-standard benchmark.

A6. A genetic-algorithm optimizer games the penalized-logP scoring function

Finding: A reproduction study of a published genetic-algorithm result found that the reported scores were reproducible mainly by exploiting deficiencies of the fitness function, which produced long, sulfur-containing chains.

Source: Jablonka K.M., Mcilwaine F., Garcia S., Smit B., Yoo B., "A reproducibility study of 'Augmenting Genetic Algorithms with Deep Neural Networks for Exploring the Chemical Space'," arXiv:2102.00700, submitted 1 February 2021, last revised 10 February 2021. Preprint, not peer-reviewed.

URL: https://arxiv.org/abs/2102.00700

Excerpt:

"Overall, we were able to reproduce comparable results using the SELFIES-based GA, but mostly by exploiting deficiencies of the (easily optimizable) fitness function (i.e., generating long, sulfur containing chains)."

Context: Another concrete, named instance of objective-gaming: a headline "state of the art" was largely an artifact of exploiting the reward. It documents reward hacking independently of neural architectures.

Confidence: Medium. Preprint reproduction study, not peer-reviewed; the finding is consistent with A2 and A3. Two longer quotations in an earlier draft could not be found in any version of the preprint and were removed.

A7. PMO benchmark: under realistic oracle budgets most molecular optimizers fail

Finding: When sample efficiency, the number of oracle evaluations and so the real constraint on experimental testing, is enforced, "most 'state-of-the-art' methods fail to outperform their predecessors under a limited oracle budget allowing 10K queries" and "no existing algorithm can efficiently solve certain molecular optimization problems in this setting."

Source: Gao W., Fu T., Sun J., Coley C.W., "Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization," NeurIPS 2022 Datasets and Benchmarks. December 2022, peer-reviewed.

URL: https://papers.nips.cc/paper_files/paper/2022/hash/8644353f7d307baaf29bc1e56fe8e0ec-Abstract-Datasets_and_Benchmarks.html

Excerpt:

"Moreover, the sample efficiency of the optimization, the number of molecules evaluated by the oracle, is rarely discussed, despite being an essential consideration for realistic discovery applications... Our results show that most 'state-of-the-art' methods fail to outperform their predecessors under a limited oracle budget allowing 10K queries and that no existing algorithm can efficiently solve certain molecular optimization problems in this setting."

Context: With a finite experimental budget, which candidates get experimentally tested is determined by the optimizer's search trajectory, and the quality of that trajectory is method-dependent and often poor. Different algorithms present different menus to the wet lab, and the wet lab cannot see what was omitted.

Confidence: High. Tier-1 venue, peer-reviewed.

A8. Over-specialization bias in growing chemical databases

Finding: Chemical databases grow in a biased way, over-specializing on studied scaffolds; models trained on known drugs fail to generalize to unexplored compounds, a selection bias; and if the target domain "cannot be specified or is generally unknown," none of the standard debiasing approaches can be applied.

Source: Dost K., Pullar-Strecker Z., Brydon L., Zhang K., Hafner J., Riddle P.J., Wicker J.S., "Combatting over-specialization bias in growing chemical databases," Journal of Cheminformatics (2023). 19 May 2023, peer-reviewed.

URL: https://jcheminf.biomedcentral.com/articles/10.1186/s13321-023-00716-w ; https://pmc.ncbi.nlm.nih.gov/articles/PMC10197453/

Excerpt:

"A bias violates this assumption and causes the generalization step to fail resulting in poor performance of the model... An example of covariate shift occurs in the drug discovery process where predictive models are trained on known drugs but expected to generalize to unexplored compounds. If the models are expected to perform well on the observed and the target data, that is, if the target space contains the observed data, this special scenario is called a Selection Bias."

"However, if the target domain cannot be specified or is generally unknown, as it is in our problem statement, none of these approaches can be used."

Context: Peer-reviewed confirmation that chemical-space coverage bias is structurally self-reinforcing, since databases grow where models already work, and that the unobserved alternatives are by construction invisible to standard correction methods. This is the "downstream validation may never reveal the omitted alternatives" question stated in the cheminformatics literature.

Confidence: High. Peer-reviewed.

A9. SimDMTA: the query strategy, not the science, determines which hits are found

Finding: In a full in silico simulation of the design, make, test, analyze cycle, the choice of acquisition strategy determines hit discovery. Under scrambled, uninformative initial training data, exploitative or greedy strategies "failed to identify any hits over 30 simulated iterations", while explorative strategies retained discovery capability, and strategies including randomness "correct systematic biases more rapidly."

Source: Williams H.J., Pickett S.D., Baxter A., Palmer D.S., "Query Matters: How Selection Strategies Influence Active Learning in Drug Discovery," J. Chem. Inf. Model. 2026, 66, 3288-3301. Peer-reviewed; GSK and Strathclyde authors.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC13014458/ ; https://strathprints.strath.ac.uk/95343/13/Williams-etal-JCIM-2026-how-selection-strategies-influence-active-learning-in-drug-discovery.pdf

Excerpt:

"Under these scrambled conditions, greedy acquisition strategies (MP and MPO) failed to identify any hits over 30 simulated iterations. In contrast, the explorative MU strategy retained some discovery capability, achieving 40% hit identification within the top 50 predicted compounds by the final iteration."

"Strategies that include some random selection correct systematic biases more rapidly, but are less effective at predicting top-performing molecules."

"To assess the potential influence of initial model bias, each query strategy was also evaluated using a control model... removing the feature-target correlation substantially impaired the model's ability to identify compounds with favorable docking scores across all iterations."

Context: The selection policy, an operational and configuration decision, determines which compounds are ever tested, and interacts badly with initial model bias. A closed design-make-test-analyze loop with a greedy policy and a biased starting model produces an empty or narrow menu, and nothing inside the loop flags the omission.

Confidence: High. Peer-reviewed, industry authors, purpose-built simulation.

A10. Bayesian optimization: surrogate misspecification biases sampling of the solution space

Finding: Bayesian optimization's search efficiency depends on surrogate quality: "the model can be sensitive to misspecification, and even slight misrepresentations of the model can introduce undesired bias, skewing the sampling of potential solutions."

Source: Liu T. et al., "Large Language Models to Enhance Bayesian Optimization," ICLR 2024 (arXiv:2402.03921). February 2024.

URL: https://arxiv.org/html/2402.03921v1

Excerpt:

"Given that BO is designated for scenarios with limited observations, constructing an accurate surrogate model with sparse observations is inherently challenging. Additionally, the model can be sensitive to misspecification, and even slight misrepresentations of the model can introduce undesired bias, skewing the sampling of potential solutions."

Context: A Tier-1-venue statement of surrogate-model misspecification leading to biased exploration. Bayesian optimization is the dominant acquisition engine in self-driving labs, see A11 and B3, so its failure mode propagates into which experiments physically happen.

Confidence: High for the mechanism. The quote is from the paper's introduction summarizing known limitations of Bayesian optimization.

A11. Closed-loop autonomous materials discovery: CAMEO, and mitigations that are configuration choices

Finding: CAMEO, a closed-loop Bayesian active-learning system at a synchrotron beamline, autonomously selects "the most informative next material to study to achieve user-defined objectives". Its own methods acknowledge "myopia to particular phase regions" as a failure mode removable only by "an additional exploration policy."

Source: Kusne A.G. et al., "On-the-fly closed-loop materials discovery via Bayesian active learning," Nature Communications (2020). Peer-reviewed.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC7686338/

Excerpt:

"Physics-informed active learning is then used to identify the most informative next material to study to achieve user-defined objectives."

"Myopia to particular phase regions can be removed with an additional exploration policy."

Context: A flagship closed-loop success story that nonetheless documents three points: the algorithm selects what is measured, so it sets the menu; the objective is user-defined, an input that can be wrong or can go out of date; and regional myopia is a known failure mode handled only by explicit anti-myopia configuration. The loop does not self-correct for menu-narrowing unless someone configures it to.

Confidence: High. Peer-reviewed, primary.

A12. The A-Lab dispute: a closed loop self-validated "novel" materials that outside scientists contested

Finding: The A-Lab, reported in Nature in November 2023, claimed the autonomous synthesis of novel materials over 17 days of continuous operation; the published critique addresses 43 claimed synthetic products. Leeman, Liu, Stiles, Lee, Bhatt, Schoop and Palgrave re-examined all of them and concluded that no new materials had been discovered, that automated Rietveld analysis of powder X-ray diffraction data "is not yet reliable", and that two thirds of the claimed successes were likely known compositionally disordered versions of the predicted ordered compounds. Nature issued an Author Correction on 19 January 2026; the paper's title now reads "accelerated synthesis of inorganic materials" and its abstract reports 36 compounds from a set of 57 targets. C&EN reported on 29 January 2026 that the critics remained unsatisfied.

Source: Szymanski N.J. et al., "An autonomous laboratory for the accelerated synthesis of inorganic materials," Nature 624, 86-91 (2023), DOI 10.1038/s41586-023-06734-w, with Author Correction DOI 10.1038/s41586-025-09992-y dated 19 January 2026; Leeman J., Liu Y., Stiles J., Lee S.B., Bhatt P., Schoop L.M., Palgrave R.G., "Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis," PRX Energy 3, 011002 (7 March 2024), peer-reviewed and open access; C&EN coverage of the correction, 29 January 2026.

URL: https://doi.org/10.1103/PRXEnergy.3.011002 ; https://doi.org/10.1038/s41586-025-09992-y ; https://cen.acs.org/research-integrity/Nature-robot-chemist-paper-corrected/104/web/2026/01

Excerpt, C&EN:

"A 2024 critique by University College London solid-state chemist Robert Palgrave and colleagues found that the materials created by A-Lab, which is housed at the Lawrence Berkeley National Laboratory, already exist in the Inorganic Crystal Structure Database, contradicting the Nature study's claims about novelty."

Excerpt, the PRX Energy critique:

"We discuss all 43 synthetic products and point out four common shortfalls in the analysis. These errors unfortunately lead to the conclusion that no new materials have been discovered in that work."

"Automated Rietveld analysis of powder x-ray diffraction data is not yet reliable."

Context: The strongest real-world case that internal validation inside a closed loop is not independent validation. The loop generated candidates, synthesized them, characterized them and pronounced them novel, and the failure was caught by outside humans re-examining the evidence. It shows both that closed-loop systems can amplify initial model and analysis bias into confidently wrong conclusions, and that the corrective came from outside the system, by scientific review.

Confidence: High. Widely reported, the Nature correction is on record, and the underlying critique is by named domain experts.

A13. Human and AI feedback loops amplify initial bias

Finding: Training AI on slightly biased human data amplifies the bias, and subsequent human interaction with the biased AI further increases human bias over time: a documented algorithmic bias feedback loop.

Source: Glickman M., Sharot T., "How human-AI feedback loops alter human perceptual, emotional and social judgements," Nature Human Behaviour (2025). Peer-reviewed.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC11860214/

Excerpt:

"These results demonstrate an algorithmic bias feedback loop; training an AI algorithm on a set of slightly biased human data results in the algorithm amplifying it. Subsequent interactions of other humans with this algorithm further increase the humans' initial bias levels, creating a feedback loop."

"...when interacting with the AI this rate increased significantly to 56.3%... The learned bias increased over time: in the first interaction block it was only 50.72%, whereas in the last interaction block it was 61.44%."

Context: General-domain but transferable: closed-loop discovery systems where model outputs determine the next training batch are a feedback loop of this kind. It grounds the claim that closed-loop systems amplify initial bias in controlled experimental evidence rather than speculation. The transfer to chemistry is an inference, stated as such.

Confidence: High for the experimental evidence.

A14. Deployed clinical models alter the data they are later judged on

Finding: Successful deployed classifiers change the distribution of features and labels, a feedback loop. Feedback-naive monitoring yields the most inaccurate performance estimates, and with true drift, standard retraining dropped AUROC from 0.72 to 0.52 while feedback-aware strategies recovered to 0.67.

Source: "Monitoring Strategies for Continuous Evaluation of Deployed Clinical Prediction Models," Journal of Biomedical Informatics (2025), 168:104854. June 2025, peer-reviewed.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC12278811/

Excerpt:

"Successful clinical machine learning classifiers will lead to a change in care which may change the distribution of features, labels, and their relationship... Classifier surveillance systems naive to such deployment-induced feedback loops will estimate lower model performance and lead to degraded future classifier retrains."

"Furthermore, in simulations with true data drift, retraining using standard unweighted approaches results in a AUROC score of 0.52 (drop from 0.72)."

Context: An adjacent domain, clinical prediction, but a rigorous demonstration that a model's own actions contaminate the evidence available for its evaluation. The remediation here is a modeling and statistics fix, feedback-aware weighting. The drug-discovery parallel is an analogy and is flagged as one.

Confidence: High for the clinical-domain result.

A15. ADMET leaderboard audit: most top-ranked models irreproducible, data leakage documented

Finding: An end-user audit of all 22 TDC ADMET leaderboards found only 3 of the top-ranked models, CaliciBoost, MapLight and MapLight+GNN, passed all checks. Direct or indirect data leakage was traced in MiniMol, GradientBoost and XGBoost, and deliberately overfitted models reached the top three in 10 of 22 endpoints against 2 of 22 for honest models.

Source: "An End-User Audit of Reproducibility, Data Leakage, and Overfitting of the Top-Ranked ADMET Prediction Models in TDC Leaderboards," J. Chem. Inf. Model. (2026), peer-reviewed, ACS. Audit performed with a leaderboard snapshot of November 2025.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC13417885/

Excerpt:

"Only three models (CaliciBoost, MapLight, and MapLight + GNN) passed all stages and reproduced their reported performance. The remaining models failed because of unavailable code, nonreproducible environments, runtime incompatibilities, or methodological flaws. We traced direct or indirect data leakage in the MiniMol, GradientBoost, and XGBoost models."

"Our deliberately overfitted models entered the top-3 in 10 out of 22 end points, whereas 'honest' models did so in only 2 out of 22. This contrast demonstrates that public leaderboards are highly susceptible to accidental or deliberate overfitting to the open testing data set."

Context: The scoring and validation layer itself can be gamed or hollow. If the triage models that decide which candidates advance are benchmark-validated in a way this easy to inflate, then validation of the menu-setter is weaker than practitioners assume. The audit had to be done by outsiders: the leaderboard itself did not surface the problems.

Confidence: High. Peer-reviewed audit.

A16. Training-data bias: selective reporting and uneven coverage of chemical space

Finding: The bioactivity database most generative models are trained on is built largely from the medicinal-chemistry literature, and carries selective reporting and uneven coverage of chemical and biological space, so the generative menu is constrained before any objective is specified.

Source: Council on Pharmacy Standards, CAIDRA examination guide, "4.2 Generative chemistry models (molecules, peptides)", industry training material, not peer-reviewed.

URL: https://pharmacystandards.org/caidra-examination/section-4-2-generative-chemistry-models-molecules-peptides/

Excerpt:

"ChEMBL is an important curated bioactivity resource built largely from medicinal-chemistry literature, while also incorporating deposited screening data and other sources. Like any aggregated dataset, it can contain missingness, assay heterogeneity, varying measurement conditions, selective reporting, duplicate chemistry, and uneven coverage of chemical/biological space."

Context: Establishes the upstream data-level bias that constrains the generative menu before any objective is specified. This is a lower-tier source, industry training material, but consistent with the peer-reviewed selection-bias literature at A8. A longer excerpt on publication bias and "dark" chemical space quoted in an earlier draft was not found on the page as served and was removed.

Confidence: Medium-high for the finding, corroborated by A8 and the QSAR bias literature; low to medium for the specific framing, since this source is not peer-reviewed.

A17. Goodhart's law, strong version: efficiently optimizing a proxy degrades the true objective

Finding: When a measure becomes a target, and if it is effectively optimized, the quantity it was designed to measure grows worse. This is presented as a general phenomenon in machine learning, independent of architecture.

Source: Sohl-Dickstein J., "Too much efficiency makes everything worse: overfitting and the strong version of Goodhart's law," technical essay, 6 November 2022.

URL: https://sohl-dickstein.github.io/2022/11/06/strong-Goodhart.html

Excerpt:

"If we keep on optimizing the proxy objective, even after our goal stops improving, something more worrying happens. The goal often starts getting worse, even as our proxy objective continues to improve... This is an extremely general phenomenon in machine learning."

Context: The theoretical umbrella over A2, A3 and A6. Not peer-reviewed, but authored by a recognized machine-learning researcher and consistent with the peer-reviewed molecular-design evidence above. Used as framing, not as primary evidence.

Confidence: Medium. Authoritative essay, not peer-reviewed; the phenomenon is independently documented in A2, A3 and A6.

A18. In practice, AI ADMET triage already gates which compounds reach an assay

Finding: Discovery organizations score full virtual libraries and advance or deprioritize compounds on model predictions before any wet-lab experiment, so models function as the gate on experimental capacity.

Source: Drug Discovery News, "AI-Powered ADMET prediction..." (2026); and the review Volkamer A., Riniker S., Nittinger E., Lanini J., Grisoni F., Evertsson E., Rodríguez-Pérez R. et al., "Machine learning for small molecule drug discovery in academia and industry," Artificial Intelligence in the Life Sciences 3, 100056 (2023), peer-reviewed.

URL: https://www.drugdiscoverynews.com/ai-powered-admet-prediction-how-machine-learning-is-changing-drug-candidate-selection-17356 ; https://doi.org/10.1016/j.ailsci.2022.100056

Excerpt, Drug Discovery News:

"Score full virtual libraries against multiple ADMET endpoints simultaneously, using predictions to rank rather than eliminate candidates outright. Route top-ranked compounds through targeted in vitro confirmation..."

Context: Model predictions already decide which physical experiments happen, and which never do, in routine industry practice. Note that the trade source states the normative ideal, rank rather than eliminate, while practice and incentives push toward elimination. A quotation attributed to the review in an earlier draft could not be checked against the text in this pass, because the copy cited was no longer reachable, and it was removed.

Confidence: High that prediction-led triage is standard practice; medium for how often it eliminates rather than reorders, which is practice-level and varies by organization.

Part II. Authority drift: AI acting under authority or constraints that no longer represent the current human decision

B1. MegaSyn: the safeguard the authors identify is an external human authority check

Finding: The only barrier that kept the inverted MegaSyn system from proceeding toward an autonomous toxin design-make-test cycle was a human decision to stop, and the authors explicitly pose the removal or replacement of that human authority as the risk scenario.

Source: Urbina et al., Nature Machine Intelligence (2022), full citation at A1. March 2022.

URL: https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/

Excerpt:

"Importantly, we had a human-in-the-loop with a firm moral and ethical 'don't-go-there' voice to intervene. But what if the human was removed or replaced with a bad actor? With current breakthroughs and research into autonomous synthesis, a complete design-make-test cycle applicable to making not only drugs, but toxins, is within reach. Our proof-of-concept highlights how a non-human autonomous creator of a deadly chemical weapon is entirely feasible."

Context: The clearest published articulation of authority drift in drug discovery. The constraint, penalize toxicity, was a sign in a reward function, changeable in minutes, and its enforcement depended entirely on a human being present, attentive and empowered. Nothing in the system itself encoded or enforced the human's authority. The change of a single reward sign is the limiting case: the system's operating constraints ceased to represent any legitimate human decision, and the only barrier was an ad hoc human veto.

Confidence: Very high. Primary, Tier-1.

B2. FDA's AI framework centres human-led governance, but discovery-stage AI sits outside the draft guidance's scope

Finding: FDA's 2023 discussion paper frames risk around human-led governance, accountability and transparency, and the peer-reviewed commentary on it notes that such systems may "amplify errors or preexisting biases in their training data". The January 2025 draft guidance proposes a risk-based credibility assessment per context of use, but it covers AI used to produce information or data that supports regulatory decision-making, and discovery-stage tools are presently out of scope, while EMA's framework is broader. FDA's own guidance page still lists it as a draft.

Source: FDA, "Using Artificial Intelligence and Machine Learning in the Development of Drug and Biological Products", discussion paper, May 2023, revised February 2025, Federal Register notice 11 May 2023; FDA, "Considerations for the Use of Artificial Intelligence to Support Regulatory Decision Making for Drug and Biological Products", draft guidance, January 2025; FDA CDER AI page; Lenarczyk G., Minssen T., Price N., Rai A., "The future of AI regulation in drug development: a comparative analysis", Journal of Law and the Biosciences 12, lsaf028 (7 November 2025), peer-reviewed.

URL: https://www.federalregister.gov/documents/2023/05/11/2023-09985/using-artificial-intelligence-and-machine-learning-in-the-development-of-drug-and-biological ; https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development ; https://pmc.ncbi.nlm.nih.gov/articles/PMC12598624/ ; https://www.drugdiscoverynews.com/fda-s-action-plan-for-ai-in-drug-development-what-scientists-need-to-know-17367

Excerpt, Lenarczyk et al. on the FDA framework:

"The FDA's 2023 discussion paper crystallized this regulatory philosophy, introducing a risk-based framework that emphasizes three interconnected domains: human-led governance with clear accountability and transparency; quality and reliability of data; and robust model development, performance, and validation protocols."

Excerpt, Lenarczyk et al. on the risk:

"...amplify errors or preexisting biases in their training data, raising questions about the generalizability of their insights across diverse patient populations."

Excerpt, FDA's CDER page on the scope of the draft guidance:

"This guidance provides recommendations to industry on the use of AI to produce information or data intended to support regulatory decision-making regarding safety, effectiveness, or quality for drugs."

Excerpt, Drug Discovery News practice note:

"Do not assume a discovery-stage tool is out of scope indefinitely; FDA's current draft excludes discovery today, but EMA's parallel framework already does not, and the two agencies are visibly converging."

Context: Regulators have named the authority question, human-led governance and accountability, but the binding machinery, credibility assessment and context-of-use statements, attaches where AI output enters a regulatory submission, far downstream of where the menu was set in discovery. The upstream search-space determination in discovery is not addressed by this framework, and no regulator requires a record of which candidates a model suppressed.

Confidence: High for the regulatory facts; medium-high for the characterization of the scope gap. The FDA draft scope is explicit; the inference about discovery being outside it follows from the scope statements. A characterization of the FDA paper quoted from a law-firm note in an earlier draft could not be verified and was replaced by the peer-reviewed wording above.

B3. Self-driving labs: the algorithm selects each successive experiment, and the objective is set once

Finding: By the field-standard definition, a self-driving lab is "a machine-learning-assisted modular experimental platform that iteratively operates a series of experiments selected by the machine learning algorithm to achieve a user-defined objective": the objective is supplied at setup and the loop then executes it. A review of autonomous labs concludes that "the absence of a global regulatory framework leaves major accountability gaps", with unauthorized or jurisdiction-shopping experimentation a live risk.

Source: Abolhasani M., Kumacheva E., "The rise of self-driving labs in chemical and materials sciences", Nature Synthesis 2, 483-492 (30 January 2023), peer-reviewed; Asghari J., Naderi M., Aboutalebi S.H., Moosavi-Movahedi A.A., "Autonomous Labs and Scientific Equity: Opportunities and Challenges", Nanoscale and Advanced Materials 2(4), 298 (1 December 2025), DOI 10.22034/nsam.2026.570277.1059. Industry survey of the field: IntuitionLabs, "Self-Driving Labs in Pharma" (2026).

URL: https://www.nature.com/articles/s44160-022-00231-0 ; https://www.jnsam.com/article_243373.html ; https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd

Excerpt, Nature Synthesis:

"...a machine-learning-assisted modular experimental platform that iteratively operates a series of experiments selected by the machine learning algorithm to achieve a user-defined objective."

Excerpts, the autonomous-labs review:

"At the same time, the absence of a global regulatory framework leaves major accountability gaps. Regulatory inconsistencies between countries create opportunities for misuse, such as remotely ordering restricted materials from permissive regions, while VPNs and decentralized access obscure identities and jurisdictions."

"Without coordinated governance, autonomous labs could become havens for unauthorized, dangerous, or illegal activities conducted outside the scope of local law."

"The presence of autonomous labs does not eliminate the need for human judgment; if anything, it amplifies its importance... ethical oversight must begin at the design stage."

Context: The technical literature itself describes the structure: the objective and constraints are inputs fixed at configuration time, and the loop then executes them autonomously. If the user who defined the objective leaves, the programme's priorities change, or the constraint file is edited, whether deliberately or by routine update, nothing in the loop architecture detects the change. Note that the industry survey of self-driving labs states that "the academic literature is more measured than vendor press releases", so the true extent of unattended operation in company platforms is not established by their public disclosures.

Confidence: High for the definitional and technical facts, now taken from the primary; medium for the accountability-gap review, which is peer-reviewed but in a lower-tier journal. The authority-drift framing is this report's synthesis and is labelled as inference. A sentence attributed to a 2026 Digital Discovery paper in an earlier draft was quoted through the industry survey and could not be checked against the journal, which blocks automated access, so it was removed.

B4. The audit-trail expectation already exists in adjacent regulated work, and AI breaks naive provenance

Finding: Regulated pharmaceutical and clinical contexts already demand reconstructable audit trails: ALCOA++, EU AI Act Article 12 logging, and the EU GMP Chapter 4 draft which adds "Traceable". Unmanaged AI use creates "loss of provenance and traceability", "undisclosed AI influence on study decisions" and "accountability gaps across sponsors and partners".

Source: Clinical Leader, "Understanding And Preserving Data Flow Integrity In AI-Assisted Clinical Trials"; Pienomial, "Building an Audit Trail for AI-Assisted HTA Submissions" (2026); Danaher Life Sciences compliance overview (2026).

URL: https://www.clinicalleader.com/doc/understanding-and-preserving-data-flow-integrity-in-ai-assisted-clinical-trials-0001 ; https://www.pienomial.com/blog/audit-trail-ai-assisted-hta-submissions ; https://lifesciences.danaher.com/us/en/library/regulatory-compliance-ai-drug-discovery.html

Excerpt, Clinical Leader:

"Loss of provenance and traceability. AI-assisted transformations may lack documented inputs, model details, version history, or reproducibility, making it difficult to explain how conclusions were reached... Undisclosed AI influence on study decisions... Accountability gaps across sponsors and partners. When multiple parties rely on AI to support judgment, it may be unclear who owns the output, who reviewed it, and who is responsible if it is challenged."

"If AI influenced a trial decision or record, can the sponsor demonstrate how it was used, by whom, in what environment, and under what review controls?"

Context: These are industry and regulatory-adjacent sources, not Tier-1, but they document the compliance demand side: the evidentiary requirement for decision provenance around AI-assisted consequential work is real and is being formalized through ALCOA++, EU AI Act Article 12 and the GMP Chapter 4 draft. These sources concern clinical and health-technology-assessment work rather than discovery.

Confidence: Medium-high. Consistent across multiple independent industry sources and anchored in real regulatory instruments, but not peer-reviewed.

B5. What was not found

Systematic searching found no documented, named case in drug discovery where an AI system was publicly shown to have continued operating under demonstrably stale or revoked human authority, for example an objective function changed without re-approval, or a retired programme's model still gating experiments, with measured consequences. No such case was found in this pass.

The evidence for authority drift is therefore of three kinds: the architectural fact that objectives and constraints are static inputs with no authority check at the moment of action (B1, B3); the regulatory demand for human-led governance and reconstructable provenance that current systems do not meet (B2, B4); and the MegaSyn demonstration that constraints are trivially changeable with no system-level barrier (B1 and A1). On the public record, authority drift in discovery-stage drug AI is a structurally demonstrated risk rather than a documented incident class.

Confidence: High that this negative result reflects the public record as searched in this pass, across multiple query strategies. Non-public industry incidents cannot be excluded.

Part III. What the literature proposes, and what it leaves open

The mitigations the cited papers themselves propose are statistical and model-internal, not organizational. Reward hacking is addressed by applicability-domain constraints and reliability-adjusted rewards (A2); mode collapse by sampling and metric choices (A4, A5); surrogate misspecification by better priors and exploration policies (A10, A11); feedback-loop bias by weighted monitoring and retraining (A13, A14); benchmark gaming by hidden test sets and versioning (A15); regional myopia by an explicit exploration policy (A11).

Two questions the literature leaves open:

  • The harm is an absence, not an event. A scientifically biased but correctly configured model narrows the menu, and the candidate that was never generated leaves no trace in the assay data. Nothing in any of the cited closed-loop architectures detects an omission while it is happening; in the A-Lab case (A12) the correction came from outside scientists re-reading the published evidence.
  • Discovery-stage configuration sits outside the regulatory frameworks. FDA's draft guidance attaches where AI output enters a regulatory submission and excludes discovery (B2); the self-driving-lab literature records the absence of a coordinated framework for autonomous labs (B3). No regulator asks for a record of which candidates a model suppressed.

Net assessment

  • The hypothesis is strongly supported for phenomenon A, scientific bias. Multiple Tier-1, named, dated cases show that objectives, acquisition policies, sampling configurations and training data determine, narrow and sometimes game the search space, with downstream validation structurally unable to reveal omissions (A1 to A18).
  • It is structurally but not yet incidentally supported for phenomenon B. The architecture allows authority drift to occur silently, and the regulatory literature asks for exactly the provenance that current systems do not produce, but no public, documented drug-discovery incident of harm from stale authority was found in this pass.

Sources

  1. Urbina F. et al. Dual use of artificial intelligence-powered drug discovery. Nat Mach Intell (2022). https://pmc.ncbi.nlm.nih.gov/articles/PMC9544280/
  2. Yoshizawa T. et al. A data-driven generative strategy to avoid reward hacking in multi-objective molecular design. Nat Commun (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC11897179/
  3. Renz P. et al. On failure modes in molecule generation and optimization. Drug Discov Today Technol 32-33, 55-63 (2019). https://doi.org/10.1016/j.ddtec.2020.09.003 ; Langevin M., Vuilleumier R., Bianciotto M. Explaining and avoiding failure modes in goal-directed generation of small molecules. J Cheminformatics (2022). https://pmc.ncbi.nlm.nih.gov/articles/PMC8973583/
  4. Özçelik R., Grisoni F. How evaluation choices distort the outcome of generative drug discovery. J Cheminformatics 17, 169 (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC12613558/ ; arXiv:2501.05457
  5. Brown N. et al. GuacaMol: Benchmarking Models for De Novo Molecular Design. J Chem Inf Model (2019). https://arxiv.org/pdf/1811.09621
  6. Jablonka K.M., Mcilwaine F., Garcia S., Smit B., Yoo B. A reproducibility study of "Augmenting Genetic Algorithms with Deep Neural Networks...". arXiv:2102.00700 (2021). https://arxiv.org/abs/2102.00700
  7. Gao W., Fu T., Sun J., Coley C.W. Sample Efficiency Matters: A Benchmark for Practical Molecular Optimization. NeurIPS (2022). https://papers.nips.cc/paper_files/paper/2022/hash/8644353f7d307baaf29bc1e56fe8e0ec-Abstract-Datasets_and_Benchmarks.html
  8. Dost K., Pullar-Strecker Z., Brydon L., Zhang K., Hafner J., Riddle P.J., Wicker J.S. Combatting over-specialization bias in growing chemical databases. J Cheminformatics (2023). https://pmc.ncbi.nlm.nih.gov/articles/PMC10197453/
  9. Williams H.J. et al. Query Matters: How Selection Strategies Influence Active Learning in Drug Discovery. J Chem Inf Model 2026, 66, 3288-3301. https://pmc.ncbi.nlm.nih.gov/articles/PMC13014458/
  10. Liu T. et al. Large Language Models to Enhance Bayesian Optimization. ICLR 2024. https://arxiv.org/html/2402.03921v1
  11. Kusne A.G. et al. On-the-fly closed-loop materials discovery via Bayesian active learning. Nat Commun (2020). https://pmc.ncbi.nlm.nih.gov/articles/PMC7686338/
  12. Szymanski N.J. et al. An autonomous laboratory for the accelerated synthesis of inorganic materials. Nature 624, 86-91 (2023), with Author Correction, 19 January 2026. https://doi.org/10.1038/s41586-025-09992-y ; Leeman J. et al. Challenges in High-Throughput Inorganic Materials Prediction and Autonomous Synthesis. PRX Energy 3, 011002 (2024). https://doi.org/10.1103/PRXEnergy.3.011002 ; C&EN, "Nature robot chemist paper corrected..." (29 January 2026). https://cen.acs.org/research-integrity/Nature-robot-chemist-paper-corrected/104/web/2026/01
  13. Glickman M., Sharot T. How human-AI feedback loops alter human perceptual, emotional and social judgements. Nat Hum Behav (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC11860214/
  14. Monitoring Strategies for Continuous Evaluation of Deployed Clinical Prediction Models. J Biomed Inform 168:104854 (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC12278811/
  15. An End-User Audit of Reproducibility, Data Leakage, and Overfitting of the Top-Ranked ADMET Prediction Models in TDC Leaderboards. J Chem Inf Model (2026). https://pmc.ncbi.nlm.nih.gov/articles/PMC13417885/
  16. Council on Pharmacy Standards, CAIDRA examination guide, "4.2 Generative chemistry models (molecules, peptides)". https://pharmacystandards.org/caidra-examination/section-4-2-generative-chemistry-models-molecules-peptides/
  17. Sohl-Dickstein J. Too much efficiency makes everything worse (6 November 2022). https://sohl-dickstein.github.io/2022/11/06/strong-Goodhart.html
  18. Drug Discovery News, "AI-Powered ADMET prediction..." (2026). https://www.drugdiscoverynews.com/ai-powered-admet-prediction-how-machine-learning-is-changing-drug-candidate-selection-17356
  19. Volkamer A., Riniker S., Nittinger E., Lanini J., Grisoni F., Evertsson E., Rodríguez-Pérez R. et al. Machine learning for small molecule drug discovery in academia and industry. Artif Intell Life Sci 3, 100056 (2023). https://doi.org/10.1016/j.ailsci.2022.100056
  20. FDA. Using AI/ML in the Development of Drug and Biological Products, discussion paper. Federal Register, 11 May 2023. https://www.federalregister.gov/documents/2023/05/11/2023-09985/using-artificial-intelligence-and-machine-learning-in-the-development-of-drug-and-biological
  21. FDA CDER. Artificial Intelligence for Drug Development, including the January 2025 draft guidance. https://www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development
  22. Lenarczyk G., Minssen T., Price N., Rai A. The future of AI regulation in drug development: a comparative analysis. J Law Biosci 12, lsaf028 (2025). https://pmc.ncbi.nlm.nih.gov/articles/PMC12598624/
  23. Drug Discovery News, "FDA's Action Plan for AI in Drug Development" (2026). https://www.drugdiscoverynews.com/fda-s-action-plan-for-ai-in-drug-development-what-scientists-need-to-know-17367
  24. Abolhasani M., Kumacheva E. The rise of self-driving labs in chemical and materials sciences. Nat Synth 2, 483-492 (2023). https://www.nature.com/articles/s44160-022-00231-0
  25. IntuitionLabs. Self-Driving Labs in Pharma: Closed-Loop R&D That Works (2026). https://intuitionlabs.ai/articles/self-driving-labs-pharma-rd
  26. Asghari J., Naderi M., Aboutalebi S.H., Moosavi-Movahedi A.A. Autonomous Labs and Scientific Equity: Opportunities and Challenges. Nanoscale and Advanced Materials 2(4), 298 (2025). https://www.jnsam.com/article_243373.html
  27. Clinical Leader. Understanding And Preserving Data Flow Integrity In AI-Assisted Clinical Trials. https://www.clinicalleader.com/doc/understanding-and-preserving-data-flow-integrity-in-ai-assisted-clinical-trials-0001
  28. Pienomial. Building an Audit Trail for AI-Assisted HTA Submissions (2026). https://www.pienomial.com/blog/audit-trail-ai-assisted-hta-submissions
  29. Danaher Life Sciences. AI in Drug Development: Regulatory Compliance Challenges (2026). https://lifesciences.danaher.com/us/en/library/regulatory-compliance-ai-drug-discovery.html