Early Prediction of At-Risk MCA Students Using Machine Learning-Based Classification Models

Authors

  • Mrs Arati Patil Research Scholar, Bharati Vidyapeeth Deemed to Be University, Pune Author
  • Dr Prashant P Patil Assistant Professor, D Y Patil Institute of MCA and Management, Akurdi, Pune Author
  • Dr S V Deshmukh Research Expert, Department of Computer Applications, Bharati Vidyapeeth Deemed to Be University, Pune Author

DOI:

https://doi.org/10.47392/IRJAEM.2026.0392

Keywords:

Educational Data Mining, Learning Analytics, Academic Performance Prediction, Machine Learning, Student Classification, At-Risk Students, Supervised Learning, Precision, Recall, F1-score

Abstract

Student academic performance prediction has become an important area of educational data mining due to its potential to improve learning outcomes through timely intervention. Early identification of academically at-risk students enables educational institutions to provide personalized academic support, thereby reducing failure rates and improving student retention. Master of Computer Applications (MCA) programmes present unique challenges because students come from diverse academic backgrounds and are expected to acquire advanced technical competencies within a limited duration. Consequently, identifying students who may experience academic difficulties at an early stage is essential for enhancing educational quality and student success. This paper proposes a machine learning-based methodology for the early prediction of MCA students' academic performance using multiclass classification. The proposed framework categorizes students into three performance groups: At-Risk, Average Performer, and High Performer. The methodology integrates student demographic information, academic history, attendance records, continuous assessment scores, laboratory performance, assignment completion, learning management system engagement, and classroom participation as predictive features. A systematic pre-processing pipeline consisting of data cleaning, feature encoding, normalization, missing value treatment, and feature selection is incorporated to improve data quality before model training. The proposed methodology recommends the implementation and comparative evaluation of several supervised machine learning algorithms, including Logistic Regression, Decision Tree, Random Forest, Support Vector Machine, K-Nearest Neighbour, Naïve Bayes, Gradient Boosting, and Extreme Gradient Boosting (XGBoost). Model performance is assessed using Accuracy, Precision, Recall, F1-score, Confusion Matrix, and Receiver Operating Characteristic (ROC) analysis. Greater emphasis is placed on Precision, Recall, and F1-score since these measures provide more reliable evaluation in educational datasets where class imbalance frequently exists. The proposed framework is expected to support higher education institutions in establishing an intelligent early warning system capable of identifying academically vulnerable students before the final examinations. Such predictive systems can facilitate timely counselling, personalized mentoring, remedial instruction, and data-driven academic decision-making, ultimately contributing to improved student retention and academic excellence.

Downloads

Download data is not yet available.

Downloads

Published

2026-10-07