Overfitting is the excessive adaptation of a machine learning model to its training data and remains a persistent challenge in biomedical informatics. In supervised learning, models may capture patterns “too well,” failing to generalize to unseen data and yielding overly optimistic performance estimates. The problem is especially acute in biomedical settings, where datasets are high-dimensional, heterogeneous, and often of few samples. Numerous strategies have been proposed to mitigate overfitting in bioinformatics and health informatics. However, even established techniques can produce misleading results if misapplied, for example through data leakage, excessive hyperparameter tuning, inappropriate preprocessing, or inadequate validation. To address these pitfalls, we present eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies. The recommendations stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis. Rather than offering an exhaustive treatment, we provide an accessible, practice-oriented guide to support more reliable and reproducible machine learning research. Although developed for biomedical informatics, these quick tips are broadly applicable across disciplines using supervised machine learning.

Chicco, D., Oneto, L. (2026). Eleven quick tips to reduce overfitting in machine learning. INTERNATIONAL JOURNAL OF DATA SCIENCE AND ANALYTICS, 22(1) [10.1007/s41060-026-01246-y].

Eleven quick tips to reduce overfitting in machine learning

Chicco D.
Primo
;
2026

Abstract

Overfitting is the excessive adaptation of a machine learning model to its training data and remains a persistent challenge in biomedical informatics. In supervised learning, models may capture patterns “too well,” failing to generalize to unseen data and yielding overly optimistic performance estimates. The problem is especially acute in biomedical settings, where datasets are high-dimensional, heterogeneous, and often of few samples. Numerous strategies have been proposed to mitigate overfitting in bioinformatics and health informatics. However, even established techniques can produce misleading results if misapplied, for example through data leakage, excessive hyperparameter tuning, inappropriate preprocessing, or inadequate validation. To address these pitfalls, we present eleven practical tips for reducing unintentional overfitting in supervised biomedical machine learning studies. The recommendations stress principled data splitting, domain-informed preprocessing, controlled model complexity, systematic tuning, comprehensive performance evaluation, and robustness analysis. Rather than offering an exhaustive treatment, we provide an accessible, practice-oriented guide to support more reliable and reproducible machine learning research. Although developed for biomedical informatics, these quick tips are broadly applicable across disciplines using supervised machine learning.
Articolo in rivista - Review Essay
Binary classification; Guidelines; Machine learning; Overfitting; Recommendations; Regression analysis; Supervised machine learning;
English
4-ago-2026
2026
22
1
256
open
Chicco, D., Oneto, L. (2026). Eleven quick tips to reduce overfitting in machine learning. INTERNATIONAL JOURNAL OF DATA SCIENCE AND ANALYTICS, 22(1) [10.1007/s41060-026-01246-y].
File in questo prodotto:
File Dimensione Formato  
Chicco-Oneto-2026-International Journal of Data Science and Analytics-VoR.pdf

accesso aperto

Tipologia di allegato: Publisher’s Version (Version of Record, VoR)
Licenza: Creative Commons
Dimensione 973.28 kB
Formato Adobe PDF
973.28 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10281/626162
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? 0
Social impact