Development of a machine learning model for loan default risk prediction in savings and credit cooperative organizations
DOI:
https://doi.org/10.51867/ajernet.7.3.141Keywords:
Credit Risk, Explainable AI, Loan Default Prediction, Machine Learning, SACCOs, SHAPAbstract
Loan default remains a substantial challenge for Savings and Credit Cooperative Organizations (SACCOs), as it can increase credit risk, reduce profitability, and weaken financial sustainability. Although SACCOs increasingly use computerized information systems to manage lending activities, historical borrower data are often not fully utilized for predictive credit-risk assessment. This study developed an explainable machine learning model for loan default risk prediction in SACCOs. The study was guided by Information Asymmetry Theory and Credit Rationing Theory, which provide a basis for understanding how borrower information can reduce uncertainty and support credit allocation decisions. Four objectives guided the study: to identify factors associated with loan default risk, develop machine learning models for loan default prediction, evaluate their predictive performance, and provide an interpretable model for credit-risk assessment. A quantitative predictive analytics design was adopted using 1,000 records from the German Credit Dataset. Data preparation involved cleaning, categorical encoding, feature scaling, feature engineering, and class balancing using the Synthetic Minority Oversampling Technique (SMOTE). Logistic Regression and Random Forest models were developed using hyperparameter optimization and Stratified 5-Fold Cross-Validation. Model performance was evaluated using Accuracy, Precision, Recall, F1-score, ROC-AUC, Brier Score, confusion-matrix analysis, and false-negative analysis. SHapley Additive exPlanations (SHAP) was subsequently used to interpret the selected model. The results indicated that checking account status, age, credit amount, and loan duration were important predictors of default risk. Logistic Regression outperformed Random Forest across all reported evaluation measures, achieving Accuracy of 77.0%, Precision of 60.6%, Recall of 66.7%, F1-score of 63.5%, ROC-AUC of 0.795, and Brier Score of 0.162. It also recorded fewer false-negative predictions, with 20 compared with 26 for Random Forest. SHAP analysis identified checking account status and age as the most influential predictors. The study concludes that Logistic Regression with SHAP-based explainability provides a practical and interpretable approach to credit-risk prediction under the conditions examined. However, the model should be validated using institution-specific SACCO data before operational implementation.
Downloads
References
Abdou, H. A., & Pointon, J. (2011). Credit scoring, statistical techniques and evaluation criteria: A review of the literature. Intelligent Systems in Accounting, Finance and Management, 18(2-3), 59-88. https://doi.org/10.1002/isaf.325
Addo, P. M., Guegan, D., & Hassani, B. (2018). Credit risk analysis using machine and deep learning models. Risks, 6(2), Article 38. https://doi.org/10.3390/risks6020038
Breiman, L. (2001). Random forests. Machine Learning, 45, 5-32. https://doi.org/10.1023/A:1010933404324
Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2020). Explainable AI in fintech risk management. Frontiers in Artificial Intelligence, 3, Article 26. https://doi.org/10.3389/frai.2020.00026
Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953
Crook, J. N., Edelman, D. B., & Thomas, L. C. (2007). Recent developments in consumer credit risk assessment. European Journal of Operational Research, 183(3), 1447-1465. https://doi.org/10.1016/j.ejor.2006.09.100
Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly Media.
James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. https://doi.org/10.1007/978-1-0716-1418-1
Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124-136. https://doi.org/10.1016/j.ejor.2015.05.030
Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in neural information processing systems 30 (pp. 4765-4774). Curran Associates, Inc.
Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML '05) (pp. 625-632). Association for Computing Machinery. https://doi.org/10.1145/1102351.1102430
Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206-215. https://doi.org/10.1038/s42256-019-0048-x
Stiglitz, J. E., & Weiss, A. (1981). Credit rationing in markets with imperfect information. American Economic Review, 71(3), 393-410.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Joseph Airo Anyimbi, Raphael Angulu, Jairus Odawa

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.













