Artificial intelligence is rapidly lowering the cost of generating plausible hypotheses in drug discovery. Models can nominate targets, integrate biological and chemical data, design molecules, and suggest experiments. Experienced discovery teams, however, have always had more ideas than time, capital, or experimental capacity to pursue them.
What is changing is the imbalance. AI can expand the number of plausible possibilities much faster than our ability to evaluate them. If evaluation capacity does not improve in parallel, cheaper generation can increase the downstream burden—more candidates, mechanistic questions, and false positives competing for finite experimental and development resources—without improving R&D productivity.
Discrimination has always been hard. The question is whether AI can help us become better at it as generation becomes increasingly abundant.
Why Crowding Persists
Large areas of biology remain comparatively underexplored, while resources repeatedly converge on targets and mechanisms supported by stronger existing evidence. That concentration is often rational. Better-established biology can reduce one dimension of risk, even when it creates intense competition for patients, trial sites, capital, and commercial differentiation.
Novel biology carries the opposite trade-off. The opportunity may be larger, but the tools, assays, mechanistic understanding, human evidence, and confidence required to evaluate it are often weaker.
AI can help expand the search by connecting genetics, multi-omics, structural biology, chemistry, and scientific literature. It can also reproduce biases embedded in those data and direct attention toward what is already well studied. The value therefore lies not simply in nominating more novel targets, but in making unfamiliar biology more defensible and testable.
Drug Discovery Is a Learning Process
Every discovery program is a sequence of hypotheses and decisions. Is the biology causally connected to disease? Can it be modulated with an appropriate therapeutic modality? Does a molecule engage the intended target with the required selectivity and exposure? Is an observed phenotype driven by the proposed mechanism? Are there safety liabilities? Will any of these findings predict what happens in patients?
Experiments generate evidence that strengthens or weakens these hypotheses. Their value depends not only on technical quality, but also on predictive validity: whether the readout is meaningfully related to the clinical outcome we ultimately care about.
A perfectly optimized loop built around a poor disease model or non-predictive assay can simply learn the wrong thing faster.
Some uncertainty can be reduced preclinically; other uncertainty can only be resolved in humans. Ultimately, efficacy and safety in patients are the arbiters of whether a drug-discovery hypothesis becomes a medicine. AI cannot remove that boundary, but it can help us arrive there with stronger evidence and fewer unresolved assumptions.
Optimize Decision-Relevant Learning
Consider a program in which a compound produces an encouraging cellular phenotype but several mechanisms could explain the result. Repeating the same assay may confirm reproducibility while adding little information about which mechanism is responsible. A more discriminating experiment could materially change the decision about whether the program should advance.
Rather than measuring an AI-enabled discovery system primarily by predictions generated, compounds designed, or experiments completed, we should also ask how much decision-relevant learning it produces per experiment, per month, and per dollar invested.
AI is not yet reliably capable of deciding which uncertainty matters most across a complex drug program. That judgment includes biology, strategy, competitive context, and risk tolerance. Where hypotheses and experimental readouts are sufficiently defined, however, AI can help rank alternatives, expose inconsistencies, estimate uncertainty, and prioritize informative experiments. Given the pace of development, the boundary of what can be delegated to computational systems is likely to continue to advance.
The objective is evidence that changes a consequential decision: advance, stop, or reframe.
Closing the Experimental Loop
The principle is already visible across therapeutic modalities. Protein and antibody design are among the clearest examples: AI can navigate vast sequence spaces and support optimization across properties such as affinity, specificity, stability, and developability before experimental testing. In RNA therapeutics, generative models have designed full-length mRNA sequences with experimentally demonstrated gains in translation and stability. And in small molecules, a 2025 prospective study used machine-learning-guided iterative screening to test only 5.9% of a two-million-compound library while recovering 43.3% of the primary actives found by a parallel full screen and nearly all compound series selected by medicinal chemists.
These examples matter because reducing the number of candidates that require physical testing can change the experiments we can afford to run. Historically, screening hundreds of thousands or millions of compounds often favored simple, highly scalable assays. If AI can narrow that search intelligently, more complex and potentially more translationally relevant systems—including patient-derived iPSC models, organoids, and other higher-content assays—can move earlier in discovery. Many of these systems are themselves becoming increasingly amenable to automation and scaled experimental workflows.
Laboratory automation can therefore contribute more than throughput. By reducing the cost and cycle time of informative experiments, it can help make richer experimental systems practical earlier. The goal should be to waste fewer experiments.
The Scarce Capability
As generative models become more capable and widely available, producing another hypothesis will become progressively less differentiating. The scarce capability will be an experimental and organizational system that challenges hypotheses rigorously, generates evidence with genuine translational relevance, and acts on what that evidence says.
A well-evidenced decision to stop a weak program can create as much value as a decision to advance a strong one. Organizations therefore need to reward learning and disciplined termination, not progression alone.
AI will not fix poor incentives, weak models, or uncertain translation by itself. Its larger opportunity is to help teams ask better questions, connect those questions to more predictive experiments, and recognize sooner when the evidence no longer supports the hypothesis.
When hypotheses become cheap, competitive advantage will not come from having the most ideas. It will come from learning, with the strongest evidence available, which ideas deserve the next experiment—and which are most likely to survive the ultimate tests of efficacy and safety in patients.

