Preprint
Machine Learning

Causability and explainability of artificial intelligence in medicine

Andreas Holzinger(Medical University of Graz), Georg Langs(Medical University of Vienna), Helmut Denk(Medical University of Graz), Kurt Zatloukal(Medical University of Graz), Heimo Müller(Medical University of Graz)
April 2, 2019Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery1,780 citations

1.8k

Citations

33

Influential Citations

Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery

Venue

2019

Year

Abstract

Explainable artificial intelligence (AI) is attracting much interest in medicine. Technically, the problem of explainability is as old as AI itself and classic AI represented comprehensible retraceable approaches. However, their weakness was in dealing with uncertainties of the real world. Through the introduction of probabilistic learning, applications became increasingly successful, but increasingly opaque. Explainable AI deals with the implementation of transparency and traceability of statistical black‐box machine learning methods, particularly deep learning (DL). We argue that there is a need to go beyond explainable AI. To reach a level of explainable medicine we need causability. In the same way that usability encompasses measurements for the quality of use, causability encompasses measurements for the quality of explanations. In this article, we provide some necessary definitions to discriminate between explainability and causability as well as a use‐case of DL interpretation and of human explanation in histopathology. The main contribution of this article is the notion of causability, which is differentiated from explainability in that causability is a property of a person, while explainability is a property of a system This article is categorized under: Fundamental Concepts of Data and Knowledge > Human Centricity and User Interaction

Analysis

Why This Paper Matters

This paper addresses a critical gap in the explainable AI (XAI) literature by shifting focus from system-centric explanations to human-centric evaluation. While much of XAI research concentrates on making black-box models transparent, Holzinger et al. argue that transparency alone is insufficient for high-stakes domains like medicine. They introduce the concept of 'causability'—a measurable property of a person's ability to derive causal understanding from explanations—paralleling the well-established concept of usability in human-computer interaction. This reframing is significant because it grounds AI explainability in cognitive science and human factors, emphasizing that the ultimate goal is not just interpretable models but effective human decision-making.

The paper's timing (2019) was prescient, as deep learning was rapidly entering clinical workflows. By distinguishing explainability (system property) from causability (human property), the authors provide a vocabulary for evaluating whether explanations actually improve clinician understanding and trust. This has influenced subsequent work on explanation quality metrics and user studies in medical AI.

Technical Contributions

  • Definition of causability: A property of a person that measures the quality of explanations, analogous to usability measuring the quality of user interfaces.
  • Clear differentiation: Explainability is a system property (how transparent a model is), while causability is a human property (how well a person can understand and reason with explanations).
  • Use-case in histopathology: Demonstrates the distinction by showing how a deep learning model's feature attribution (explainability) differs from a pathologist's causal reasoning (causability) when interpreting tissue slides.
  • Framework for evaluation: Proposes that causability can be measured through user studies, similar to usability testing, opening the door for empirical validation of explanation methods.

Results

The paper does not present quantitative results or benchmarks. Its primary contribution is conceptual: establishing the need for causability as a distinct evaluation dimension. The use-case illustrates the gap between model explanations (e.g., saliency maps) and human causal understanding, but no metrics are provided.

Significance

This paper has been highly influential (1780 citations) in shaping the discourse on explainable AI, particularly in medicine. It has motivated a shift from purely technical XAI research toward human-centered evaluation, inspiring work on explanation quality metrics, user studies, and regulatory guidelines for AI in healthcare. The concept of causability has been adopted in subsequent frameworks for trustworthy AI and remains relevant as AI systems become more integrated into clinical decision-making.