Anomaly Detection Feature Engineering for Enhanced Prediction of HIV Viral Load Suppression Using DHIS2 Data
Jonathan Ndolo Mbithi
Department of Mathematics and Statistics, University of Embu, Embu, Kenya.
Maurice Wanyonyi *
Department of Mathematics and Statistics, University of Embu, Embu, Kenya.
Rodgers Mavyala
Department of Computer Science, Presbyterian University of East Africa, Kikuyu, Kenya.
Collins Odhiambo
Institute of Mathematical Sciences, Strathmore University, Nairobi, Kenya.
*Author to whom correspondence should be addressed.
Abstract
Highly imbalanced, incomplete, and complex feature datasets such as routine HIV programme data in DHIS2 are increasingly used for predictive modelling. Yet conventional feature engineering often fails to capture subtle, non-linear patterns that can enable accurate model predictions. This study aimed to evaluate whether integrating unsupervised anomaly detection into feature engineering improves the prediction of HIV viral load suppression from routine DHIS2 data compared with conventional feature-engineering approaches. A phased analytical pipeline compared raw-variable evaluation using 63,594 patient records from two Kenyan counties. Conventional domain-driven feature engineering and anomaly-detection-based feature engineering were compared using Isolation Forest, Local Outlier Factor, and PCA reconstruction-error metrics. Retention-in-care variables consistently demonstrated the strongest associations with appointment-related metrics, showing effect sizes of approximately 0.99. Conventional engineering improved interpretability but did not substantially enhance predictive performance. Anomaly detection identified 5–10% of patients as anomalous, and twelve anomaly-aware features were generated. Five of these, including lof_score and if_score, were ranked among the top 20 predictors, and their inclusion increased XGBoost discrimination from an AUC of 0.801 to 0.823 (p<0.001). Cross-phase analysis confirmed that anomaly-derived features complemented rather than displaced established clinical predictors. The findings indicate that anomaly detection can serve as an effective feature-engineering strategy, offering a scalable approach to strengthening predictive modelling using routinely collected, complex health data in resource-constrained settings.
Keywords: Anomaly detection, feature engineering, HIV viral load suppression, DHIS2, Machine learning, isolation forest, local outlier factor, principal component analysis, XGBoost, routine health data