From Explainable to Contestable AI: Lessons from Medical AI Ethics for Drug Discovery

Challenging AI

In drug discovery, an AI-generated target ranking can influence experiments, portfolio choices and capital allocation long before a candidate enters clinical development. The question is therefore not only whether the model can explain its output, but whether the organization can challenge and, where warranted, revise the decision it informs.

Imagine a research meeting. An AI system has prioritized a therapeutic target. Its predictions look promising, and an explanation highlights the features behind its ranking. The team is considering the next steps. Then someone asks: Does this really make sense? The important question is what happens next. Does the concern prompt further investigation, or does the discussion move on because the model has provided an apparently sufficient explanation? And who decides what counts as sufficient?

My research on medical AI ethics and explainability has led me to ask how explanations can support judgment, responsibility, participation, and learning in healthcare. Healthcare delivery and drug discovery differ in their purposes, risks, and evidentiary demands. Yet the ethical starting point is the same. AI’s potential is too important to dismiss, but too consequential to leave unquestioned.

Explainable AI has become an established area of drug-discovery research. Yet making a prediction more transparent does not necessarily enable a team to challenge the decision it informs. This matters because AI does more than accelerate analysis, it gradually shapes professional roles and expectations, forms of collaboration, and our ideas of what counts as valid evidence.

Explanation, Evidence, and Judgment

An account of why a model ranked a target highly, an explanation of a biological mechanism, and a justification for investing in that target answer different questions. Moving between them requires additional evidence, judgment, and ethical reflection. In healthcare, core principles such as autonomy, non-maleficence, beneficence, and justice remind us that technically convincing recommendations still have to be weighed against human interests, potential harms, benefits, and fairness. In drug discovery, the equivalent question is not only whether a model is accurate, but what kind of decisions it is helping to make and who may be affected by those decisions downstream.

In our work on clinical decision support, colleagues and I examined how explainability requirements depend on the technology, its users, and its role in decision-making. An explanation that helps a developer investigate model behavior may leave a doctor’s, nurse’s, or patient’s question unanswered or may simply be incomprehensible to them. A plausible explanation may also fail to capture the actual underlying process or encourage an unwarranted causal interpretation.

For discovery teams, this means asking which claims an explanation supports and which uncertainties remain. The evidence needed to justify a limited experiment differs from the evidence needed to commit substantial resources or abandon an alternative. Teams therefore need to ask what would justify the next step, and what result would make them reconsider.

This is not to romanticize human judgment or position it as superior. Its essential role is to preserve forms of doubt and accountability on which responsible decisions depend. A model does not worry that it may have missed something – a team can. The capacity to pause, ask whether the available evidence is enough, and remain accountable for the next step is an inherent part of scientific integrity.

Making AI-Supported Decisions Contestable

Return to the meeting. A biologist questions whether the experimental context supports the inference. A data scientist asks whether the model generalizes beyond its training conditions. The R&D lead questions whether the evidence warrants the proposed investment. Each concern draws on expertise and experience that the others may lack. Interdisciplinary collaboration becomes meaningful when those concerns receive a response. Who investigates them? Can they lead to another experiment, a revised assumption or a different decision?

These questions connect to the concept of contestable AI, which in simple terms refers to systems that remain open to human intervention. In practice, contestability means that an AI-supported recommendation has a named decision owner, documented assumptions and uncertainties, a route for domain experts to raise concerns, and a defined mechanism to investigate and resolve them before the next commitment is made. Some of the most relevant competencies in this context include recognizing uncertainty, understanding the scope of a model’s claims, and knowing when another discipline’s expertise is needed. Teams need enough openness to explore what AI can reveal, and enough critical distance to ask whether dissent can be voiced without being treated as resistance to progress.

Challenging Both AI and Human Judgment

Human judgment is not neutral and must also remain open to challenge. Evidence shows how implicit biases can shape how healthcare professionals interpret symptoms, assess risks, and recommend treatments, contributing to differences in care and outcomes between patient groups. Similarly, in a study of compound prioritization, medicinal chemists showed limited agreement and were not always aware of the criteria influencing their selections. Although the study concerned a specific task, it cautions against treating human intuition as an unquestionable standard.

When their conclusions diverge, teams should examine the evidence, assumptions and limitations on both sides. Contestability creates the conditions for doing so before confidence in either the model or the expert closes down further inquiry.

Including patients in questions of purpose

An important question that technical explanations cannot settle is “Why is this the objective worth pursuing?” Here, patient participation becomes relevant to research agenda-setting and the interface between discovery and development. Patients can contribute knowledge of unmet needs, treatment burdens, and what meaningful improvement would look like. Their contribution concerns the purposes research serves and should not require a deep understanding of the science and technology behind it.

Healthcare has already shown how biased or incomplete datasets can reproduce unequal access and systematically disadvantage groups that are less well represented. For drug discovery, this raises questions about whose biology, symptoms, treatment burdens, and unmet needs become visible to a model, and whose remain outside the frame. Yet being present in a dataset is not the same as having a voice in how research priorities are defined. If AI is used to infer patient needs from narratives, organizations should still ask whether patients can contest those interpretations and whether their input can change the agenda.

From Principles to Practice

In January 2026, the EMA and FDA published their “Guiding Principles of Good AI Practice in Drug Development”, recognizing the importance of multidisciplinary expertise, human-AI interaction and clear information for intended audiences. The practical challenge is to translate these principles into routines and processes that affect decision-making.

For discovery teams, one useful exercise is to trace a recent AI-supported decision. What did the model contribute? What remained uncertain? Who raised a concern, how was it addressed, and why did the team proceed or change course? This exercise could reveal where an explanation helped, where additional expertise was needed, or where a concern had no clear route to resolution.

Realizing the full potential of AI in drug discovery will require more than better explanations. R&D leaders must create clear routes for challenging AI-supported decisions and the human judgments surrounding them, investigating concerns and revising course when the evidence demands it.

References

  1. Alfrink, K., Keller, I., Kortuem, G., & Doorn, N. (2023). Contestable AI by design: Towards a framework. Minds and Machines, 33, 613–639.
  2. Amann, J., Blasimme, A., Vayena, E., Frey, D., & Madai, V. I. (2020). Explainability for artificial intelligence in healthcare: A multidisciplinary perspective. BMC Medical Informatics and Decision Making, 20(1), Article 310.
  3. Amann, J., Bürger, V. K., Livne, M., Bui, C. K., & Madai, V. I. (2025). The fundamentals of AI ethics in medical imaging. In Trustworthy AI in Medical Imaging(pp. 7-33). Academic Press.
  4. Amann, J., Vetter, D., Blomberg, S. N., Christensen, H. C., Coffee, M., Gerke, S., … & Z-Inspection Initiative. (2022). To explain or not to explain?—Artificial intelligence explainability in clinical decision support systems. PLOS Digital Health, 1(2), e0000016.
  5. Bürger, V. K., Amann, J., Bui, C. K. T., Fehr, J., & Madai, V. I. (2024). The unmet promise of trustworthy AI in healthcare: Why we fail at clinical translation. Frontiers in Digital Health, 6, Article 1279629.
  6. European Medicines Agency & U.S. Food and Drug Administration. (2026, January). Guiding principles of good AI practice in drug development. URL: https://www.ema.europa.eu/en/documents/other/guiding-principles-good-ai-practice-drug-development_en.pdf (Accessed 11.09.2026)
  7. Hall, W. J., Chapman, M. V., Lee, K. M., Merino, Y. M., Thomas, T. W., Payne, B. K., Eng, E., Day, S. H., & Coyne-Beasley, T. (2015). Implicit racial/ethnic bias among health care professionals and its influence on health care outcomes: A systematic review. American Journal of Public Health, 105(12), e60–e76.
  8. Jiménez-Luna, J., Grisoni, F., & Schneider, G. (2020). Drug discovery with explainable artificial intelligence. Nature Machine Intelligence, 2, 573–584.
  9. Kutchukian, P. S., Vasilyeva, N. Y., Xu, J., Lindvall, M. K., Dillon, M. P., Glick, M., … & Brooijmans, N. (2012). Inside the mind of a medicinal chemist: the role of human bias in compound prioritization during drug discovery. PloS one, 7(11), e48476.
Views expressed are my own and do not reflect those of the Careum Foundation.

Hot Topics

Related Articles