Skip to content
Data Science & information systems International Journal of Advances in Data and Information Systems
Open access E-ISSN 2721-3056 Acceptance rate: 28%

AugLog-LightGBM: A Log-Based Feature AugmentationFramework for Class Imbalance in Credit RiskClassification

Authors

  • Hana Azizah Department of Statistics, Brawijaya University, Malang, Indonesia image/svg+xml
  • Eni Sumarminingsih Department of Statistics, Brawijaya University, Malang, Indonesia image/svg+xml
  • Adji Achmad Rinaldo Fernandes Department of Statistics, Brawijaya University, Malang, Indonesia image/svg+xml

DOI:

https://doi.org/10.59395/ijadis.v7i2.1623

Keywords:

LightGBM, log-based feature augmentation, logit, log-density ratio, class imbalance

Abstract

Non-performing loan (NPL) detection is inherently a class-imbalance problem because defaulting borrowers represent a persistent minority. Standard gradient boosting often favors the majority class. This paper proposes AugLog-LightGBM, an extension of LightGBM that improves initialization through Log-Based Feature Augmentation (LBFA). Instead of using an uninformative constant, boosting starts from an informed prior combining a logistic-regression logit score and a kernel-density-estimation log-density ratio (LDR), which capture complementary global and local information. These representations are incorporated as augmented features and as the init_score, reformulating boosting as residual correction over an informed Bayesian prior. The proposed framework is evaluated on a dataset of 2,700 home-mortgage borrowers collected from partner banks in Malang, Indonesia (NPL rate = 16.11%), using repeated stratified cross-validation and comparison against four imbalance-aware baselines. AugLog-LightGBM achieves the highest ROC-AUC (0.815 ± 0.018), PR-AUC (0.679), and F1-score (0.631). DeLong tests show statistically significant ROC-AUC improvements over class-weighted Logistic Regression and Random Forest, while gains over XGBoost and SMOTE + LightGBM are positive but not statistically significant. Robustness analyses and SHAP interpretation further support the consistency and practical applicability of the proposed framework.

481 107

Downloads

Download data is not yet available.

References

[1] Otoritas Jasa Keuangan, "Siaran Pers RDKB Desember 2024," Jakarta, Indonesia, 2024. [Online]. Available: https://www.ojk.go.id

[2] Otoritas Jasa Keuangan, "Peraturan OJK No. 15/POJK.03/2017 tentang Penetapan Status dan Tindak Lanjut Pengawasan Bank Umum," Jakarta, Indonesia, 2017.

[3] H. He and E. A. Garcia, "Learning from imbalanced data," IEEE Trans. Knowl. Data Eng., vol. 21, no. 9, pp. 1263–1284, 2009, doi: 10.1109/TKDE.2008.239. DOI: https://doi.org/10.1109/TKDE.2008.239

[4] R. Rahmadini and B. J. Santoso, "Machine learning-based prediction of divorce verdicts using posita data and imbalanced data handling: A case study in Padang Sidempuan," Int. J. Adv. Data Inf. Syst., vol. 6, no. 2, pp. 460–478, 2025, doi: 10.59395/ijadis.v6i2.1405. DOI: https://doi.org/10.59395/ijadis.v6i2.1405

[5] M. Wati, A. Thobirin, and S. Surono, "Application of EfficientNetV2-S architecture with focal loss to overcome class imbalances in skin cancer classification," Int. J. Adv. Data Inf. Syst., vol. 7, no. 1, pp. 440–452, 2026, doi: 10.59395/ijadis.v7i1.1524. DOI: https://doi.org/10.59395/ijadis.v7i1.1524

[6] G. Ke et al., "LightGBM: A highly efficient gradient boosting decision tree," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 3146–3154.

[7] K. Pyar, "Improvement the accuracy of convolutional neural network with using undersampling method on unbalanced credit card dataset," Int. J. Adv. Data Inf. Syst., vol. 5, no. 2, pp. 183–188, 2024, doi: 10.59395/ijadis.v5i2.1333. DOI: https://doi.org/10.59395/ijadis.v5i2.1333

[8] U. Rawat and B. Rawat, "A comprehensive study of boosting algorithms for class imbalance dataset," in Smart Innovation, Systems and Technologies, 2025, pp. 269–282, doi: 10.1007/978-981-96-1348-9_21. DOI: https://doi.org/10.1007/978-981-96-1348-9_21

[9] S. Rizky, P. R. Hirzi, and H. H. Umam, "Perbandingan metode LightGBM dan XGBoost dalam menangani data dengan kelas tidak seimbang," J. Statistika, vol. 15, no. 2, pp. 228–236, 2022, doi: 10.36456/jstat.vol15.no2.a5548. DOI: https://doi.org/10.36456/jstat.vol15.no2.a5548

[10] J. Fan, Y. Feng, J. Jiang, and X. Tong, "Feature augmentation via nonparametrics and selection (FANS) in high-dimensional classification," J. Amer. Statist. Assoc., vol. 111, no. 513, pp. 275–287, 2016, doi: 10.1080/01621459.2015.1005212. DOI: https://doi.org/10.1080/01621459.2015.1005212

[11] J. Li and W. Cui, "A new classifier for imbalanced data based on a generalized density ratio model," Commun. Math. Stat., vol. 11, no. 2, pp. 369–401, 2023, doi: 10.1007/s40304-021-00254-7. DOI: https://doi.org/10.1007/s40304-021-00254-7

[12] S. Seniaray and R. Jindal, "Performance analysis of anomaly-based network intrusion detection using feature selection and machine learning techniques," Wirel. Pers. Commun., vol. 138, no. 4, pp. 2321–2351, 2024, doi: 10.1007/s11277-024-11602-5. DOI: https://doi.org/10.1007/s11277-024-11602-5

[13] J. I. Daoud, "Multicollinearity and regression analysis," J. Phys. Conf. Ser., vol. 949, no. 1, p. 012009, 2018, doi: 10.1088/1742-6596/949/1/012009. DOI: https://doi.org/10.1088/1742-6596/949/1/012009

[14] D. W. Hosmer and S. Lemeshow, Applied Logistic Regression, 2nd ed. Hoboken, NJ, USA: Wiley, 2000, doi: 10.1002/0471722146. DOI: https://doi.org/10.1002/0471722146

[15] A. Agresti, Categorical Data Analysis, 2nd ed. Hoboken, NJ, USA: Wiley, 2002, doi: 10.1002/0471249688. DOI: https://doi.org/10.1002/0471249688

[16] A. Chairunissa, S. Solimun, and A. A. R. Fernandes, "Creditor classification logistic regression ensemble boosting and logistic regression in creditor classification with binary response," WSEAS Trans. Syst. Control, vol. 16, pp. 705–714, 2021, doi: 10.37394/23203.2021.16.64. DOI: https://doi.org/10.37394/23203.2021.16.64

[17] B. W. Silverman, Density Estimation for Statistics and Data Analysis. London, U.K.: Chapman & Hall, 1986, doi: 10.1007/978-1-4899-3324-9. DOI: https://doi.org/10.1007/978-1-4899-3324-9

[18] D. W. Scott, Multivariate Density Estimation: Theory, Practice, and Visualization. New York, NY, USA: Wiley, 1992. DOI: https://doi.org/10.1002/9780470316849

[19] M. Sugiyama, T. Suzuki, and T. Kanamori, Density Ratio Estimation in Machine Learning. Cambridge, U.K.: Cambridge Univ. Press, 2012, doi: 10.1017/CBO9781139035613. DOI: https://doi.org/10.1017/CBO9781139035613

[20] J. H. Friedman, "Greedy function approximation: A gradient boosting machine," Ann. Statist., vol. 29, no. 5, pp. 1189–1232, 2001, doi: 10.1214/aos/1013203451. DOI: https://doi.org/10.1214/aos/1013203451

[21] N. A. Sovia, N. W. S. Wardhani, and E. Sumarminingsih, "Hybrid CNN-SVM with borderline SMOTE for imbalance class cabbage plants," Inferensi, vol. 7, no. 3, pp. 199–205, 2024, doi: 10.12962/j27213862.v7i3.20514. DOI: https://doi.org/10.12962/j27213862.v7i3.20514

[22] D. M. W. Powers, "Evaluation: From precision, recall and F-measure to ROC, informedness, markedness and correlation," J. Mach. Learn. Technol., vol. 2, no. 1, pp. 37–63, 2011.

[23] S. M. Lundberg et al., "From local explanations to global understanding with explainable AI for trees," Nat. Mach. Intell., vol. 2, no. 1, pp. 56–67, 2020, doi: 10.1038/s42256-019-0138-9. DOI: https://doi.org/10.1038/s42256-019-0138-9

[24] S. M. Lundberg and S.-I. Lee, "A unified approach to interpreting model predictions," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 30, 2017, pp. 4765–4774.

[25] D. J. Hand and W. E. Henley, "Statistical classification methods in consumer credit scoring: A review," J. R. Stat. Soc. Ser. A, vol. 160, no. 3, pp. 523–541, 1997, doi: 10.1111/j.1467-985X.1997.00078.x. DOI: https://doi.org/10.1111/j.1467-985X.1997.00078.x

[26] T. Chen and C. Guestrin, "XGBoost: A scalable tree boosting system," in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., 2016, pp. 785–794, doi: 10.1145/2939672.2939785. DOI: https://doi.org/10.1145/2939672.2939785

[27] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, "SMOTE: Synthetic minority over-sampling technique," J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002, doi: 10.1613/jair.953. DOI: https://doi.org/10.1613/jair.953

[28] A. Carriero, K. Luijken, A. de Hond, K. G. M. Moons, B. van Calster, and M. van Smeden, "The harms of class imbalance corrections for machine learning based prediction models: A simulation study," Stat. Med., vol. 44, no. 3–4, 2025, doi: 10.1002/sim.10320. DOI: https://doi.org/10.1002/sim.10320

[29] B. Baesens, V. Van Vlasselaer, and W. Verbeke, Fraud Analytics Using Descriptive, Predictive, and Social Network Techniques. Hoboken, NJ, USA: Wiley, 2015, doi: 10.1002/9781119146841. DOI: https://doi.org/10.1002/9781119146841

[30] E. R. DeLong, D. M. DeLong, and D. L. Clarke-Pearson, "Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach," Biometrics, vol. 44, no. 3, pp. 837–845, 1988, doi: 10.2307/2531595. DOI: https://doi.org/10.2307/2531595

Downloads

Published

2026-08-06

How to Cite

[1]
H. Azizah, E. . Sumarminingsih, and A. A. R. . Fernandes, “AugLog-LightGBM: A Log-Based Feature AugmentationFramework for Class Imbalance in Credit RiskClassification”, International Journal of Advances in Data and Information Systems, vol. 7, no. 2, pp. 772–785, Aug. 2026, doi: 10.59395/ijadis.v7i2.1623.

Share



Plum Analytics


Similar Articles

1-10 of 149

You may also start an advanced similarity search for this article.