Skip to content
Data Science & information systems International Journal of Advances in Data and Information Systems
Open access E-ISSN 2721-3056 Acceptance rate: 28%

Robust Clustering Analysis of Village Potential Data in South Sumatra Using MMCD-Based Mahalanobis Distance

Authors

DOI:

https://doi.org/10.59395/ijadis.v7i2.1568

Keywords:

Robust Mahalanobis Distance, MMCD, Clustering Analysis , Village Potential Data, Outliers

Abstract

Village potential data often contain outliers that can distort clustering results when classical distance measures are applied. This study evaluates the effectiveness of the Robust Mahalanobis Distance based on the Matrix Minimum Covariance Determinant (RMD-MMCD) for clustering village potential data in South Sumatra Province, Indonesia. The proposed approach is compared across K-means, K-medoids, DBSCAN, and DB-Kmeans algorithms, with performance assessed using the Silhouette coefficient, Dunn index, Davies–Bouldin index, Rand Index, and Adjusted Rand Index. The findings demonstrate that incorporating robust covariance estimation through MMCD improves clustering stability and enhances resistance to outliers while preserving the inherent matrix structure of multivariate data. Notable differences in cluster composition are observed primarily in the accessibility and transportation and public service dimensions. These results confirm the advantage of robust distance measures for producing more reliable and structurally consistent clustering solutions in village potential data.

413 102

Downloads

Download data is not yet available.

References

[1] N. Berliana Indah Pratiwi, Indahwati, and A. Fitrianto, Village Potential Mapping: Comprehensive Cluster Analysis of Continuous and Categorical Variables with Missing Values and Outliers Dataset in Bogor, West Java, Indonesia, Sci. J. Informatics, vol. 11, no. 2, pp. 353366, 2024, doi: 10.15294/sji.v11i2.3903. DOI: https://doi.org/10.15294/sji.v11i2.3903

[2] E. Zahid, J. Shabbir, and O. A. Alamri, A Generalized Class of Estimators for Sensitive Variable in the Presence of Measurement Error and Non-Response under Stratified Random Sampling, J. King Saud Univ. - Sci., vol. 34, no. 2, 2022, doi: 10.1016/j.jksus.2021.101741. DOI: https://doi.org/10.1016/j.jksus.2021.101741

[3] Badan Pusat Statistik, Statistik Potensi Desa Indonesia 2024, Jakarta, 2024. doi: 1105014.

[4] BPS Provinsi Sumatera Selatan, Statistik Potensi Desa Provinsi Sumatera Selatan 2024, Palembang, 2024. doi: 1105014.16.

[5] T. Li, B. yang Li, X. wei Xin, Y. yuan Ma, and Q. Yang, A Novel Tree Structure-Based Multi-Prototype Clustering Algorithm, J. King Saud Univ. - Comput. Inf. Sci., vol. 36, no. 3, p. 102002, 2024, doi: 10.1016/j.jksuci.2024.102002. DOI: https://doi.org/10.1016/j.jksuci.2024.102002

[6] X. Tao et al., Density Peak Clustering using Global and Local Consistency Adjustable Manifold Distance, Inf. Sci. (Ny)., vol. 577, pp. 769804, 2021, doi: 10.1016/j.ins.2021.08.036. DOI: https://doi.org/10.1016/j.ins.2021.08.036

[7] F. M. Barus and S. Sutarman, Mendeteksi Outlier pada Data Multivariat dengan Metode Jarak Mahalanobis-Minimum Covariance Determinant (MMCD), IJM Indones. J. Multidiscip., vol. 1, no. 3, pp. 11641172, 2023, [Online]. Available: https://journal.csspublishing/index.php/ijm

[8] R. Lakshmi and T. A. Sajesh, A Robust Distance-Based Approach for Detecting Multidimensional Outliers, J. Appl. Stat., vol. 52, no. 6, pp. 12781298, 2025, doi: https://doi.org/10.1080/02664763.2024.2422403 A. DOI: https://doi.org/10.1080/02664763.2024.2422403

[9] M. Mayrhofer, U. Radojii, P. Filzmoser, M. Mayrhofer, and U. Radoji, Robust Covariance Estimation and Explainable Outlier Detection for Matrix-Valued Data Robust Covariance Estimation and Explainable Outlier Detection for Matrix-Valued Data, Technometrics, vol. 67, pp. 561530, 2025, doi: 10.1080/00401706.2025.2475781. DOI: https://doi.org/10.1080/00401706.2025.2475781

[10] BPS Indoensia, Indeks Pembangunan Desa 2018, Jakarta, 2018. doi: 1105023.

[11] H. S. Wong and A. Fitrianto, Split Sample Sequential Fences based on Bootstrap Cut Off Points for Identifying Outliers and Parameter Estimations, ASM Sci. J., vol. 17, no. 3, 2022, doi: 10.32802/asmscj.2022.500. DOI: https://doi.org/10.32802/asmscj.2022.500

[12] H. Liu, J. Li, Y. Wu, and Y. Fu, Clustering with Outlier Removal, IEEE Trans. Knowl. Data Eng., vol. 33, no. 6, pp. 23692379, 2021, doi: 10.1109/TKDE.2019.2954317. DOI: https://doi.org/10.1109/TKDE.2019.2954317

[13] W. H. Finch, A Comparison of Clustering Methods when Group Sizes are Unequal, Outliers are Present, and in the Presence of Noise Variables, Gen. Linear Model J., vol. 45, no. 1, pp. 1222, 2019, doi: 10.31523/glmj.045001.003. DOI: https://doi.org/10.31523/glmj.045001.003

[14] A. Papadimitriou and V. Tsoukala, Evaluating and Enhancing the Performance of the K-Means Clustering Algorithm for Annual Coastal Bed Evolution Applications, Oceanologia, no. xxxx, 2024, doi: 10.1016/j.oceano.2023.12.005. DOI: https://doi.org/10.1016/j.oceano.2023.12.005

[15] J. Medellu and E. S. Nugraha, K-Medoids Clustering Applications for High-Dimensionality Multiphase, in Proceeding of The Symposium on Data Science 2021 K-MEANS, President University Actuarial Science, 2021, pp. 113. [Online]. Available: https://e-journal.president.ac.id/index.php/SDS/article/view/1726/965

[16] A. Sobrinho Campolina Martins, L. Ramos de Araujo, and D. Rosana Ribeiro Penido, K-Medoids Clustering Applications for High-Dimensionality Multiphase Probabilistic Power Flow, Int. J. Electr. Power Energy Syst., vol. 157, no. August 2023, p. 109861, 2023, doi: 10.1016/j.ijepes.2024.109861. DOI: https://doi.org/10.1016/j.ijepes.2024.109861

[17] N. F. Fahrudin and R. Rindiyani, Comparison of K-Medoids and K-Means Algorithms in Segmenting Customers based on RFM Criteria, in E3S Web of Conferences 484 (FoITIC 2023), EDP Sciences, 2024, pp. 117. doi: 10.1051/e3sconf/202448402008. DOI: https://doi.org/10.1051/e3sconf/202448402008

[18] S. S. Li, An Improved DBSCAN Algorithm Based on the Neighbor Similarity and Fast Nearest Neighbor Query, IEEE Access, vol. 8, pp. 4746847476, 2020, doi: 10.1109/ACCESS.2020.2972034. DOI: https://doi.org/10.1109/ACCESS.2020.2972034

[19] A. Latifi-Pakdehi and N. Daneshpour, DBHC: A DBSCAN-based Hierarchical Clustering Algorithm, Data Knowl. Eng., vol. 135, no. April, p. 101922, 2021, doi: 10.1016/j.datak.2021.101922. DOI: https://doi.org/10.1016/j.datak.2021.101922

[20] X. Gao, X. Ding, T. Han, and Y. Kang, Analysis of Influencing Factors on Excellent Teachers Professional Growth based on DB-Kmeans Method, EURASIP J. Adv. Signal Process., vol. 2022, no. 1, 2022, doi: 10.1186/s13634-022-00948-2. DOI: https://doi.org/10.1186/s13634-022-00948-2

[21] G. Dong, Y. Jin, S. Wang, W. Li, Z. Tao, and S. Guo, DB-Kmeans:An Intrusion Detection Algorithm Based on DBSCAN and K-means, 2019 20th Asia-Pacific Netw. Oper. Manag. Symp. Manag. a Cyber-Physical World, APNOMS 2019, pp. 14, 2019, doi: 10.23919/APNOMS.2019.8892910. DOI: https://doi.org/10.23919/APNOMS.2019.8892910

[22] P. R. Underhill and T. W. Krause, Robust Statistics and Cluster Analysis in NDT For Multi-Parameter Signals, Res. Sq., pp. 121, 2022, doi: https://doi.org/10.21203/rs.3.rs-1785606/v1. DOI: https://doi.org/10.21203/rs.3.rs-1785606/v1

[23] M. Rafiq, Y. S. Chauhan, and S. Sahay, Efficient Implementation of Mahalanobis Distance on Ferroelectric FinFET Crossbar for Outlier Detection, IEEE J. Electron Devices Soc., vol. 12, no. May, pp. 516524, 2024, doi: 10.1109/JEDS.2024.3416441. DOI: https://doi.org/10.1109/JEDS.2024.3416441

[24] . Yorulmaz, S. K. Yldrm, and B. F. Yldrm, Robust Mahalanobis Distance based Topsis to Evaluate the Economic Development of Provinces, Oper. Res. Eng. Sci. Theory Appl., vol. 4, no. 2, pp. 102123, 2021, doi: 10.31181/oresta20402102y. DOI: https://doi.org/10.31181/oresta20402102y

[25] M. Hubert and M. Debruyne, Minimum Covariance Determinant, WIREs Comput. Stat., vol. 2, no. MCD, pp. 3643, 2010, doi: https://doi.org/10.1002/wics.61. DOI: https://doi.org/10.1002/wics.61

[26] M. Mughnyanti, S. Efendi, and M. Zarlis, Analysis of Determining Centroid Clustering X-means Algorithm with Davies-Bouldin Index Evaluation, IOP Conf. Ser. Mater. Sci. Eng., vol. 725, no. 1, 2020, doi: 10.1088/1757-899X/725/1/012128. DOI: https://doi.org/10.1088/1757-899X/725/1/012128

[27] I. T. Umagapi, B. Umaternate, S. Komputer, P. Pasca Sarjana Universitas Handayani, B. Kepegawaian Daerah Kabupaten Pulau Morotai, and B. Riset dan Inovasi, Uji Kinerja K-Means Clustering Menggunakan Davies-Bouldin Index Pada Pengelompokan Data Prestasi Siswa, Semin. Nas. SISFOTEK, pp. 303308, 2023.

[28] L. Morales and J. Aguilar, An Automatic Merge Technique to Improve the Clustering Quality Performed by LAMDA, IEEE Access, vol. 8, pp. 162917162944, 2020, doi: 10.1109/ACCESS.2020.3021675. DOI: https://doi.org/10.1109/ACCESS.2020.3021675

[29] J. M. Luna-Romera, M. Martnez-Ballesteros, J. Garca-Gutirrez, and J. C. Riquelme, External Clustering Validity Index Based on Chi-Squared Statistical Test, Inf. Sci. (Ny)., vol. 487, pp. 117, 2019, doi: 10.1016/j.ins.2019.02.046. DOI: https://doi.org/10.1016/j.ins.2019.02.046

[30] S. Dogru and V. Altuntas, Prediction of Cancer in DNA Sequences using Unsupervised Learning Methods, J. Innov. Sci. Eng., vol. 7, no. 1, pp. 4047, 2022, doi: 10.38088/jise.1134816. DOI: https://doi.org/10.38088/jise.1134816

[31] D. Kallberg, L. Vidman, and P. Ryden, Comparison of Methods for Feature Selection in Clustering of High-Dimensional RNA-Sequencing Data to Identify Cancer Subtypes, Front. Genet., vol. 12, no. February, 2021, doi: 10.3389/fgene.2021.632620. DOI: https://doi.org/10.3389/fgene.2021.632620

[32] Y. Zhu, D. X. Zhang, X. F. Zhang, M. Yi, L. Ou-Yang, and M. Wu, EC-PGMGR: Ensemble Clustering Based on Probability Graphical Model with Graph Regularization for Single-Cell RNA-seq Data, Front. Genet., vol. 11, no. November, pp. 112, 2020, doi: 10.3389/fgene.2020.572242. DOI: https://doi.org/10.3389/fgene.2020.572242

[33] K. R. Shahapure and C. Nicholas, Cluster Quality Analysis using Silhouette Score, Proc. - 2020 IEEE 7th Int. Conf. Data Sci. Adv. Anal. DSAA 2020, pp. 747748, 2020, doi: 10.1109/DSAA49011.2020.00096. DOI: https://doi.org/10.1109/DSAA49011.2020.00096

[34] Y. Januzaj, E. Beqiri, and A. Luma, Determining the Optimal Number of Clusters using Silhouette Score as a Data Mining Technique, Int. J. online Biomed. Eng., vol. 19, no. 4, pp. 174182, 2023, doi: 10.3991/ijoe.v19i04.37059. DOI: https://doi.org/10.3991/ijoe.v19i04.37059

[35] N. Puspitasari, G. Lempas, H. Hamdani, H. Haviuddin, and A. Septiarini, Perbandingan Algoritma K-Means dan Algoritma K-Medoids Pada Kasus Covid-19 di Indonesia, Build. Informatics, Technol. Sci., vol. 4, no. 4, pp. 20152027, 2023, doi: 10.47065/bits.v4i4.2994. DOI: https://doi.org/10.47065/bits.v4i4.2994

[36] H. Malikhatin, A. Rusgiyono, and D. A. Maruddani, Penerapan K-Modes Clustering dengan Validasi Dunn Index pada Pengelompokan Karakteristik Calon TKI Menggunakan R-GUI, J. GAUSSIAN, vol. 10, no. 3, pp. 359366, 2021, [Online]. Available: https://ejournal3.undip.ac.id/index.php/gaussian/ DOI: https://doi.org/10.14710/j.gauss.v10i3.32790

[37] D. Chicco, A. Campagner, A. Spagnolo, D. Ciucci, and G. Jurman, The Silhouette Coefficient and the Davies-Bouldin Index are More Informative than Dunn index , Calinski-Harabasz Index , Shannon Entropy , and Gap Statistic For Unsupervised Clustering Internal Evaluation of two Convex Clusters, PeerJ Comput. Sci., pp. 149, 2025, doi: 10.7717/peerj-cs.3309. DOI: https://doi.org/10.7717/peerj-cs.3309

[38] A. M. Ikotun, F. Habyarimana, and A. E. Ezugwu, Cluster Validity Indices for Automatic Clustering: A Comprehensive Review, Heliyon, vol. 11, no. 2, p. e41953, 2025, doi: 10.1016/j.heliyon.2025.e41953. DOI: https://doi.org/10.1016/j.heliyon.2025.e41953

[39] C. Wongoutong, The Impact of Neglecting Feature Scaling in K-means Clustering, PLoS One, vol. 19, no. 12, pp. 111, 2024, doi: https://doi.org/10.1371/journal.pone.0310839. DOI: https://doi.org/10.1371/journal.pone.0310839

[40] M. Shutaywi and N. N. Kachouie, Silhouette Analysis for Performance Evaluation in Machine, Entropy, vol. 23, pp. 117, 2021, doi: https://doi.org/10.3390/e23060759. DOI: https://doi.org/10.3390/e23060759

Downloads

Published

2026-08-08

How to Cite

[1]
D. Fitrianti, A. Fitrianto, and A. Kurnia, “Robust Clustering Analysis of Village Potential Data in South Sumatra Using MMCD-Based Mahalanobis Distance”, International Journal of Advances in Data and Information Systems, vol. 7, no. 2, pp. 801–815, Aug. 2026, doi: 10.59395/ijadis.v7i2.1568.

Share



Plum Analytics


Similar Articles

1-10 of 169

You may also start an advanced similarity search for this article.