Development of a machine learning model for loan default risk prediction in savings and credit cooperative organizations

Authors

  • Joseph Airo Anyimbi Masinde Muliro University of Science and Technology, Kenya
  • Raphael Angulu Masinde Muliro University of Science and Technology, Kenya
  • Jairus Odawa Masinde Muliro University of Science and Technology, Kenya

DOI:

https://doi.org/10.51867/ajernet.7.3.141

Keywords:

Credit Risk, Explainable AI, Loan Default Prediction, Machine Learning, SACCOs, SHAP

Abstract

Loan default remains a substantial challenge for Savings and Credit Cooperative Organizations (SACCOs), as it can increase credit risk, reduce profitability, and weaken financial sustainability. Although SACCOs increasingly use computerized information systems to manage lending activities, historical borrower data are often not fully utilized for predictive credit-risk assessment. This study developed an explainable machine learning model for loan default risk prediction in SACCOs. The study was guided by Information Asymmetry Theory and Credit Rationing Theory, which provide a basis for understanding how borrower information can reduce uncertainty and support credit allocation decisions. Four objectives guided the study: to identify factors associated with loan default risk, develop machine learning models for loan default prediction, evaluate their predictive performance, and provide an interpretable model for credit-risk assessment. A quantitative predictive analytics design was adopted using 1,000 records from the German Credit Dataset. Data preparation involved cleaning, categorical encoding, feature scaling, feature engineering, and class balancing using the Synthetic Minority Oversampling Technique (SMOTE). Logistic Regression and Random Forest models were developed using hyperparameter optimization and Stratified 5-Fold Cross-Validation. Model performance was evaluated using Accuracy, Precision, Recall, F1-score, ROC-AUC, Brier Score, confusion-matrix analysis, and false-negative analysis. SHapley Additive exPlanations (SHAP) was subsequently used to interpret the selected model. The results indicated that checking account status, age, credit amount, and loan duration were important predictors of default risk. Logistic Regression outperformed Random Forest across all reported evaluation measures, achieving Accuracy of 77.0%, Precision of 60.6%, Recall of 66.7%, F1-score of 63.5%, ROC-AUC of 0.795, and Brier Score of 0.162. It also recorded fewer false-negative predictions, with 20 compared with 26 for Random Forest. SHAP analysis identified checking account status and age as the most influential predictors. The study concludes that Logistic Regression with SHAP-based explainability provides a practical and interpretable approach to credit-risk prediction under the conditions examined. However, the model should be validated using institution-specific SACCO data before operational implementation.

Downloads

Download data is not yet available.

References

Abdou, H. A., & Pointon, J. (2011). Credit scoring, statistical techniques and evaluation criteria: A review of the literature. Intelligent Systems in Accounting, Finance and Management, 18(2-3), 59-88. https://doi.org/10.1002/isaf.325

Addo, P. M., Guegan, D., & Hassani, B. (2018). Credit risk analysis using machine and deep learning models. Risks, 6(2), Article 38. https://doi.org/10.3390/risks6020038

Breiman, L. (2001). Random forests. Machine Learning, 45, 5-32. https://doi.org/10.1023/A:1010933404324

Bussmann, N., Giudici, P., Marinelli, D., & Papenbrock, J. (2020). Explainable AI in fintech risk management. Frontiers in Artificial Intelligence, 3, Article 26. https://doi.org/10.3389/frai.2020.00026

Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321-357. https://doi.org/10.1613/jair.953

Crook, J. N., Edelman, D. B., & Thomas, L. C. (2007). Recent developments in consumer credit risk assessment. European Journal of Operational Research, 183(3), 1447-1465. https://doi.org/10.1016/j.ejor.2006.09.100

Géron, A. (2022). Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow (3rd ed.). O'Reilly Media.

James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: With applications in R (2nd ed.). Springer. https://doi.org/10.1007/978-1-0716-1418-1

Lessmann, S., Baesens, B., Seow, H.-V., & Thomas, L. C. (2015). Benchmarking state-of-the-art classification algorithms for credit scoring: An update of research. European Journal of Operational Research, 247(1), 124-136. https://doi.org/10.1016/j.ejor.2015.05.030

Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in neural information processing systems 30 (pp. 4765-4774). Curran Associates, Inc.

Niculescu-Mizil, A., & Caruana, R. (2005). Predicting good probabilities with supervised learning. In Proceedings of the 22nd International Conference on Machine Learning (ICML '05) (pp. 625-632). Association for Computing Machinery. https://doi.org/10.1145/1102351.1102430

Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206-215. https://doi.org/10.1038/s42256-019-0048-x

Stiglitz, J. E., & Weiss, A. (1981). Credit rationing in markets with imperfect information. American Economic Review, 71(3), 393-410.

Downloads

Published

2026-09-21

How to Cite

Anyimbi, J. A., Angulu, R., & Odawa, J. (2026). Development of a machine learning model for loan default risk prediction in savings and credit cooperative organizations. African Journal of Empirical Research, 7(3), 1818-1827. https://doi.org/10.51867/ajernet.7.3.141