Calibration and Uncertainty Quantification for Machine Learning Models in Imbalanced HIV Viral Load Prediction: A Reliability-Based Evaluation
Jonathan Ndolo Mbithi
Department of Mathematics and Statistics, University of Embu, Embu, Kenya.
Maurice Wanyonyi *
Department of Mathematics and Statistics, University of Embu, Embu, Kenya.
Nyakora Bichang'a Simion
Department of Mathematics and Statistics, University of Embu, Embu, Kenya.
*Author to whom correspondence should be addressed.
Abstract
Machine learning (ML) models for HIV prediction are often evaluated using discrimination metrics (e.g., AUC) while neglecting calibration and predictive uncertainty, a critical gap in imbalanced clinical datasets where inaccurate probability estimates can harm decision-making. This study evaluated the calibration, uncertainty, and reliability of Logistic Regression, Random Forest, SVM, and ANN models for HIV viral load prediction using empirical DHIS2 data and a synthetic dataset. Models achieved moderate discrimination (AUC: 0.72–0.74) but very poor minority-class recall (approximately 0.6%–6.2%). Class imbalance correction markedly improved recall for several models, exceeding 60% for Logistic Regression and 78% for Linear SVM, but generally degraded calibration. Feature engineering yielded only marginal gains, whereas post-hoc calibration (Platt Scaling, Isotonic Regression) substantially improved probability reliability without sacrificing discrimination. Uncertainty and subgroup analyses revealed that strong overall performance can mask reliability deficits in clinically vulnerable groups. These results indicate that discrimination alone overestimates model trustworthiness. A comprehensive evaluation of calibration, uncertainty, and subgroup reliability is essential for the safe deployment of ML-based decision support in HIV care.
Keywords: Calibration analysis, HIV Prediction, machine learning, predictive uncertainty, probabilistic reliability