AugLog-LightGBM: A Log-Based Feature AugmentationFramework for Class Imbalance in Credit RiskClassification
DOI:
https://doi.org/10.59395/ijadis.v7i2.1623Keywords:
LightGBM, log-based feature augmentation, logit, log-density ratio, class imbalanceAbstract
Non-performing loan (NPL) detection is inherently a class-imbalance problem because defaulting borrowers represent a persistent minority. Standard gradient boosting often favors the majority class. This paper proposes AugLog-LightGBM, an extension of LightGBM that improves initialization through Log-Based Feature Augmentation (LBFA). Instead of using an uninformative constant, boosting starts from an informed prior combining a logistic-regression logit score and a kernel-density-estimation log-density ratio (LDR), which capture complementary global and local information. These representations are incorporated as augmented features and as the init_score, reformulating boosting as residual correction over an informed Bayesian prior. The proposed framework is evaluated on a dataset of 2,700 home-mortgage borrowers collected from partner banks in Malang, Indonesia (NPL rate = 16.11%), using repeated stratified cross-validation and comparison against four imbalance-aware baselines. AugLog-LightGBM achieves the highest ROC-AUC (0.815 ± 0.018), PR-AUC (0.679), and F1-score (0.631). DeLong tests show statistically significant ROC-AUC improvements over class-weighted Logistic Regression and Random Forest, while gains over XGBoost and SMOTE + LightGBM are positive but not statistically significant. Robustness analyses and SHAP interpretation further support the consistency and practical applicability of the proposed framework.
Downloads
References
[1] Otoritas Jasa Keuangan, "Siaran Pers RDKB Desember 2024," Jakarta, Indonesia, 2024. [Online]. Available: https://www.ojk.go.id
[2] Otoritas Jasa Keuangan, "Peraturan OJK No. 15/POJK.03/2017 tentang Penetapan Status dan Tindak Lanjut Pengawasan Bank Umum," Jakarta, Indonesia, 2017.
[3] H. He and E. A. Garcia, "Learning from imbalanced data," IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, 2009, doi: 10.1109/TKDE.2008.239. DOI: https://doi.org/10.1109/TKDE.2008.239
[4] R. Rahmadini and B. J. Santoso, "Machine learning-based prediction of divorce verdicts using posita data and imbalanced data handling: A case study in Padang Sidempuan," Int. J. Adv. Data Inf. Syst., vol. 6, no. 2, pp. 460–478, 2025, doi: 10.59395/ijadis.v6i2.1405. DOI: https://doi.org/10.59395/ijadis.v6i2.1405
[5] M. Wati, A. Thobirin, and S. Surono, "Application of EfficientNetV2-S architecture with focal loss to overcome class imbalances in skin cancer classification," Int. J. Adv. Data Inf. Syst., vol. 7, no. 1, pp. 440–452, 2026, doi: 10.59395/ijadis.v7i1.1524. DOI: https://doi.org/10.59395/ijadis.v7i1.1524
[6] G. Ke et al., "LightGBM: A highly efficient gradient boosting decision tree," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 3146–3154.
[7] K. Pyar, "Improvement the accuracy of convolutional neural network with using undersampling method on unbalanced credit card dataset," Int. J. Adv. Data Inf. Syst., vol. 5, no. 2, pp. 183–188, 2024, doi: 10.59395/ijadis.v5i2.1333. DOI: https://doi.org/10.59395/ijadis.v5i2.1333
[8] U. Rawat and B. Rawat, "A comprehensive study of boosting algorithms for class imbalance dataset," in Smart Innovation, Systems and Technologies, 2025, pp. 269–282, doi: 10.1007/978-981-96-1348-9_21. DOI: https://doi.org/10.1007/978-981-96-1348-9_21
[9] S. Rizky, P. R. Hirzi, and H. H. Umam, "Perbandingan metode LightGBM dan XGBoost dalam menangani data dengan kelas tidak seimbang," J. Statistika, vol. 15, no. 2, pp. 228–236, 2022, doi: 10.36456/jstat.vol15.no2.a5548. DOI: https://doi.org/10.36456/jstat.vol15.no2.a5548
[10] J. Fan, Y. Feng, J. Jiang, and X. Tong, "Feature augmentation via nonparametrics and selection (FANS) in high-dimensional classification," J. Amer. Statist. Assoc., vol. 111, no. 513, pp. 275–287, 2016, doi: 10.1080/01621459.2015.1005212. DOI: https://doi.org/10.1080/01621459.2015.1005212
[11] J. Li and W. Cui, "A new classifier for imbalanced data based on a generalized density ratio model," Commun. Math. Stat., vol. 11, no. 2, pp. 369–401, 2023, doi: 10.1007/s40304-021-00254-7. DOI: https://doi.org/10.1007/s40304-021-00254-7
[12] S. Seniaray and R. Jindal, "Performance analysis of anomaly-based network intrusion detection using feature selection and machine learning techniques," Wirel. Pers. Commun., vol. 138, no. 4, pp. 2321–2351, 2024, doi: 10.1007/s11277-024-11602-5. DOI: https://doi.org/10.1007/s11277-024-11602-5
[13] J. I. Daoud, "Multicollinearity and regression analysis," J. Phys. Conf. Ser., vol. 949, no. 1, p. 012009, 2018, doi: 10.1088/1742-6596/949/1/012009. DOI: https://doi.org/10.1088/1742-6596/949/1/012009
[14] D. W. Hosmer and S. Lemeshow, Applied Logistic Regression, 2nd ed. Hoboken, NJ, USA: Wiley, 2000, doi: 10.1002/0471722146. DOI: https://doi.org/10.1002/0471722146
[15] A. Agresti, Categorical Data Analysis, 2nd ed. Hoboken, NJ, USA: Wiley, 2002, doi: 10.1002/0471249688. DOI: https://doi.org/10.1002/0471249688
[16] A. Chairunissa, S. Solimun, and A. A. R. Fernandes, "Creditor classification logistic regression ensemble boosting and logistic regression in creditor classification with binary response," WSEAS Trans. Syst. Control, vol. 16, pp. 705–714, 2021, doi: 10.37394/23203.2021.16.64. DOI: https://doi.org/10.37394/23203.2021.16.64
[17] B. W. Silverman, Density Estimation for Statistics and Data Analysis. London, U.K.: Chapman & Hall, 1986, doi: 10.1007/978-1-4899-3324-9. DOI: https://doi.org/10.1007/978-1-4899-3324-9
[18] D. W. Scott, Multivariate Density Estimation: Theory, Practice, and Visualization. New York, NY, USA: Wiley, 1992. DOI: https://doi.org/10.1002/9780470316849
[19] M. Sugiyama, T. Suzuki, and T. Kanamori, Density Ratio Estimation in Machine Learning. Cambridge, U.K.: Cambridge Univ. Press, 2012, doi: 10.1017/CBO9781139035613. DOI: https://doi.org/10.1017/CBO9781139035613
[20] J. H. Friedman, "Greedy function approximation: A gradient boosting machine," Ann. Statist., vol. 29, no. 5, pp. 1189–1232, 2001, doi: 10.1214/aos/1013203451. DOI: https://doi.org/10.1214/aos/1013203451
[21] N. A. Sovia, N. W. S. Wardhani, and E. Sumarminingsih, "Hybrid CNN-SVM with borderline SMOTE for imbalance class cabbage plants," Inferensi, vol. 7, no. 3, pp. 199–205, 2024, doi: 10.12962/j27213862.v7i3.20514. DOI: https://doi.org/10.12962/j27213862.v7i3.20514
[22] D. M. W. Powers, "Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation," J. Mach. Learn. Technol., vol. 2, no. 1, pp. 37–63, 2011.
[23] S. M. Lundberg et al., "From local explanations to global understanding with explainable AI for trees," Nat. Mach. Intell., vol. 2, no. 1, pp. 56–67, 2020, doi: 10.1038/s42256-019-0138-9. DOI: https://doi.org/10.1038/s42256-019-0138-9
[24] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 4765–4774.
[25] D. J. Hand and W. E. Henley, "Statistical classification methods in consumer credit scoring: A review," J. R. Stat. Soc. Ser. A, vol. 160, no. 3, pp. 523–541, 1997, doi: 10.1111/j.1467-985X.1997.00078.x. DOI: https://doi.org/10.1111/j.1467-985X.1997.00078.x
[26] T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., 2016, pp. 785–794, doi: 10.1145/2939672.2939785. DOI: https://doi.org/10.1145/2939672.2939785
[27] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002, doi: 10.1613/jair.953. DOI: https://doi.org/10.1613/jair.953
[28] A. Carriero, K. Luijken, A. de Hond, K. G. M. Moons, B. van Calster, and M. van Smeden, "The harms of class imbalance corrections for machine learning based prediction models: A simulation study," Stat. Med., vol. 44, no. 3–4, 2025, doi: 10.1002/sim.10320. DOI: https://doi.org/10.1002/sim.10320
[29] B. Baesens, V. Van Vlasselaer, and W. Verbeke, Fraud Analytics Using Descriptive, Predictive, and Social Network Techniques. Hoboken, NJ, USA: Wiley, 2015, doi: 10.1002/9781119146841. DOI: https://doi.org/10.1002/9781119146841
[30] E. R. DeLong, D. M. DeLong, and D. L. Clarke-Pearson, "Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach," Biometrics, vol. 44, no. 3, pp. 837–845, 1988, doi: 10.2307/2531595. DOI: https://doi.org/10.2307/2531595
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hana Azizah, Eni Sumarminingsih, Adji Achmad Rinaldo Fernandes

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
How to Cite
Share
Plum Analytics