Comparative Performance of Expectation-Maximization (EM) Algorithm Variants for Finite Mixtures of Linear Regression Models

Hafsa Mohammed

Department of Statistics, Faculty of Science, University of Benghazi, Benghazi, Libya.

Abdelbaset Abdalla *

Department of Statistics, Faculty of Science, University of Benghazi, Benghazi, Libya.

Omar A. El-Sharif

Department of Statistics, Faculty of Science, University of Benghazi, Benghazi, Libya.

Abir Hassan Elgihawi

Department of Statistics, Faculty of Science, University of Benghazi, Benghazi, Libya.

Ahmed Mami

Department of Statistics, Faculty of Science, University of Benghazi, Benghazi, Libya.

*Author to whom correspondence should be addressed.


Abstract

Aims/ Objectives: Finite mixture models, particularly mixtures of linear regressions, are widely used for modeling heterogeneous data. Although the Expectation-Maximization (EM) algorithm is the standard estimation method, variants such as Classification EM (CEM) and Stochastic EM (SEM) offer distinct advantages. However, choosing the optimal algorithm under varying data conditions remains a practical challenge. This research addresses this gap through a systematic empirical comparison of EM, CEM, and SEM using controlled simulations that vary sample size, model complexity, mixing proportions, regression configuration, and cross-validation folds. Prediction accuracy was evaluated using the Mean Root Squared Error of Prediction (MRSEP). Our results demonstrate that sample size, model complexity, mixing proportion, and algorithm choice all significantly influence model performance, and that no single algorithm dominates across all settings. For two-component models, EM and CEM perform similarly and clearly outperform SEM, while the concurrent configuration yields substantially lower prediction errors than the parallel configuration. For three-component models, SEM is best at balanced proportions, whereas at extreme proportions EM is best at upper extremes and SEM at lower extremes for small samples. For four-component models, EM is generally most accurate at lower proportions, with CEM matching at higher proportions and SEM showing slightly higher error. Larger sample sizes generally improve accuracy, though gains vary by algorithm and model complexity. Model estimation is substantially easier at extreme mixing proportions compared to balanced proportions. The choice of cross-validation fold number has no significant impact on accuracy, confirming robustness to the evaluation procedure. Practically, these findings recommend tailoring algorithm selection to the number of components, the mixing proportion, and the sample size, rather than relying on a single universally superior method, thereby offering the scientific community an evidence-based foundation for more reliable and reproducible mixture regression modeling.

Keywords: Finite mixture models, EM Algorithm, classification EM, stochastic EM, algorithm selection, cross-validation


How to Cite

Mohammed, Hafsa, Abdelbaset Abdalla, Omar A. El-Sharif, Abir Hassan Elgihawi, and Ahmed Mami. 2026. “Comparative Performance of Expectation-Maximization (EM) Algorithm Variants for Finite Mixtures of Linear Regression Models”. Asian Journal of Probability and Statistics 28 (10):129-54. https://doi.org/10.9734/ajpas/2026/v28i10960.

Downloads

Download data is not yet available.