Approximately $10B is being invested globally in the development and application of AI in Drug Discovery in 2026. As a systems modeler, I see current investment focusing on dealing with the complicated aspects of drug discovery but not necessarily the complex aspects. Understanding this distinction is critical to developing drugs that can improve patient care and also derive commercial value.
Efforts to select potential drug targets using multi-omic data, predicting 3-D structures for these targets, generating libraries of (bio)molecular candidates, incorporating RWE and the vast published (and unpublished, proprietary) literature are being well-served using machine learning, agentic searches, LLM’s, conventional knowledge graphs and related AI methodologies. Time required to generate drug candidates is being greatly reduced and several are advancing into clinical trials. In general, however, these approaches are dealing with “known knowns” and “known unknowns” and not positioned to address the “unknown unknowns” that reflect the true complexities of the patient, of the disease and of the practice of medicine. The potential to significantly impact drug discovery requires resource allocation, in parallel efforts, to address these complexities.
- Disease is a process, not a state: we lack detailed knowledge of detailed disease trajectories as, ethically, we treat a patient when symptoms appear. Biomarkers/diagnostics reflect attempts to available technologies and are “applied” with disease risk/diagnosis/staging/ but also may vary over the course of the disease, e.g her2/neu levels. Knowledge of disease trajectories could support prevention and enhance drug development while optimizing treatment decisions
- Diagnosing/stratifying diseases and syndromes: many diseases are actually syndromes, i.e. an umbrella diagnosis including several subtypes, although using single diagnostic codes, e.g. ICD-10. Limitations in current diagnoses (see above) also occur where symptoms and clinical signs may overlap at certain stages of different diseases. This impacts treatment decisions, confounds target selection and clinical trial cohort recruitment. Examples include: hypertension, sepsis, etc
- Qualifying Real World Data and Evidence: Real world data and evidence is of increasing interest to support clinical trials and drug discovery/development and frequently is integrated into big data repositories for use with AI/ML methods. Such analysis, typically using machine learning, is used to identify complex patterns and relationships but these are correlative not causal. In addition, the umbrella of “big data” may obscure the reality that such data is truly “fit for purpose”. One of the largest segments of data, claims data, reflects business and operations, i.e. reimbursement, adherence to standard of care, etc, rather than the underlying physiology/biology. Clinical notes also require additional assessment because of the introduction of “cut and paste” into commercial EHR systems that can limit the value of data captured in these notes.
- Data Interoperability and Equivalence: Major efforts to aggregate large data sets for application of AI/ML methods also need to consider the limitations of current ontologies and EHR’s to provide adequate resolution to component data fields. Although such data fields may bear the same “name”, significant differences may exist between institutions as to how that data is generated, e.g. different laboratory test and/or evaluation of thresholds and difference in the practice of medicine and/or adherence to clinical guidelines. By example: a patient who is diagnosed as triple negative breast cancer in one institution may not receive the same diagnosis at separate institution because of differences in laboratory tests as well as thresholds that are established for +/-. This is even further exemplified by efforts to establish additional levels of her2/neu amplification (Her2(+) to Her2(-)) and, even more recently, progesterone receptor levels.
- Pathways and Drug Targets: A fundamental perspective on biology/physiology focuses on biological pathways, e.g. the intermediary metabolism “Wall Chart”. These pathways, however, represent a reductionist model of the underlying biology as 1) they are not “units” specifically coded in the genome and 2) essentially all enzymes in these pathways are promiscuous in that they can and do interact with many other molecules that are not included in the pathways as we represent them. Pathways also are rarely, if ever, linear and incorporate complex branching to effect adaptability to stresses, e.g. inhibitors, etc, to still enable their function, even if at limited capacity. If pathways were linear, then a potential “cut” across such a reaction step could prove lethal to the organism and not recoverable from the insult. This impacts the selection of targets for drug development as well as the ability to interpret overall impact of variants in any single pathway component. Example: in hypertension, ACE inhibitors are used to prevent Angiotensin I to II conversion but other enzymes, Chymase, Cathepsin G and CMLP are also capable of this conversion.
- Real world patients: Clinical trials focus cohort recruitment on establishing safety and efficacy of a drug with a bias towards expanded prescribing opportunities in the clinic. Real world patients present, however, with co-morbidities, clinical history, lifestyle and environmental impact, behavioral and cultural influences, compliance issues, physician preferences, etc all of which can be confounders in the potential for translating clinical trial success into commercial success.
- Clinical Guidelines, Standard of Care and Clinical Practice: Clinical guidelines commonly represent the evidence-based consensus of professional clinical organizations where individual steps may have varying degrees of confidence. Clinical practice involves the implementation, whole or in part, of these guidelines within the boundaries of a physician’s experience among other external considerations of the patient and the healthcare environment. The Standard of care reflects the accepted level of implementation and helps define the expected level of implementation of the guidelines. Variation in guidelines may result from multiple agencies, e.g. more than 16 organizations issue guidelines for hypertension where these separate guidelines are not harmonized in terms of specifics or periodicity of revision. Resulting variability in diagnosis can readily impact drug target selection and cohort recruitment for clinical trials.
It is critical as we explore the use of AI in drug discovery that we do not solely focus on the functions that currently appear complicated…at the expense of expanding our opportunity to explore that which is complex.

