Machine learning (ML) models are increasingly proposed to support clinical decision-making, yet their evidentiary basis remains weaker than their publication volume suggests. This editorial argues that the problem is not only translational, regulatory, or infrastructural, but methodological. Many medical ML pipelines rely on uncertain ground truths, optimize performance around clinically irrelevant thresholds, report unstable or prevalence-dependent metrics, neglect calibration and uncertainty, and lack rigorous external and temporal validation. These weaknesses produce optimistic estimates that do not reliably anticipate performance in heterogeneous clinical settings. We call for an evidence-based medical AI grounded in more reliable annotation practices, explicit modeling of uncertainty, clinically meaningful threshold selection, calibration and decision-utility analyses, robustness testing, external validation on independent datasets, and post-deployment monitoring. The editorial also invites authors, reviewers, users, and vendors to adopt stricter standards so that predictive models can become credible, accountable, and clinically useful tools in everyday practice, rather than merely publishable artifacts.

Why almost all ML models for medicine are wrong-and what we need for evidence-based medical AI

Federico Cabitza;Giuseppe Jurman;
2026-01-01

Abstract

Machine learning (ML) models are increasingly proposed to support clinical decision-making, yet their evidentiary basis remains weaker than their publication volume suggests. This editorial argues that the problem is not only translational, regulatory, or infrastructural, but methodological. Many medical ML pipelines rely on uncertain ground truths, optimize performance around clinically irrelevant thresholds, report unstable or prevalence-dependent metrics, neglect calibration and uncertainty, and lack rigorous external and temporal validation. These weaknesses produce optimistic estimates that do not reliably anticipate performance in heterogeneous clinical settings. We call for an evidence-based medical AI grounded in more reliable annotation practices, explicit modeling of uncertainty, clinically meaningful threshold selection, calibration and decision-utility analyses, robustness testing, external validation on independent datasets, and post-deployment monitoring. The editorial also invites authors, reviewers, users, and vendors to adopt stricter standards so that predictive models can become credible, accountable, and clinically useful tools in everyday practice, rather than merely publishable artifacts.
File in questo prodotto:
File Dimensione Formato  
1-s2.0-S1386505626002789-main.pdf

accesso aperto

Licenza: Creative commons
Dimensione 1.07 MB
Formato Adobe PDF
1.07 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11582/373268
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
social impact