Skip to content
Data Science & information systems International Journal of Advances in Data and Information Systems
Open access E-ISSN 2721-3056 Acceptance rate: 28%

An Integrated Data Cleaning and Anomaly Detection Pipeline for Improving Clinical Data Quality in Primary Healthcare

Authors

DOI:

https://doi.org/10.59395/ijadis.v7i2.1485

Keywords:

Clinical Data Quality, Primary Healthcare Analytics, ICD-10 Standardization, Anomaly Detection Pipeline, Data Cleaning and Preprocessing

Abstract

Clinical data quality in primary healthcare settings often faces challenges such as inconsistent terminology, missing values, duplicate entries, and relational inconsistencies across clinical variables. These issues hinder the accuracy and reliability of epidemiological analyses and predictive modeling based on real-world clinical data. This study develops and evaluates a comprehensive data-cleaning pipeline using 2,354 visit-level clinical records obtained from a primary healthcare facility, representing 496 unique patients. The proposed pipeline integrates diagnosis normalization through ICD-10 mapping supported by fuzzy matching and rule-based refinement, missing-value handling, patient identity deduplication, diagnosis–therapy relational consistency evaluation, and anomaly detection using statistical approaches (IQR, Z-score, and quantile analysis) as well as clustering techniques. The results demonstrate that 319 unique raw diagnosis strings were consolidated into 19 standardized ICD-10 codes, missing values across key clinical variables were reduced by more than 70%, and relational inconsistencies between diagnoses and therapies were substantially minimized. Visualization of diagnosis–therapy relationships reveals more coherent and interpretable clinical patterns after standardization. The novelty of this study lies in the development of a multi-method, integrated cleaning framework tailored to primary healthcare data in Indonesia. The resulting dataset exhibits improved structural integrity, terminology consistency, and relational coherence, making it more reliable and suitable for downstream analytics and predictive modeling. These findings highlight that a structured data-cleaning pipeline can significantly enhance the interpretability and analytical readiness of real-world clinical data in primary care environments.

598 146

Downloads

Download data is not yet available.

References

[1] D. J. Albers, N. Elhadad, J. Claassen, R. Perotte, A. Goldstein, and G. Hripcsak, “Estimating summary statistics for electronic health record laboratory data for use in high-throughput phenotyping algorithms,” J. Biomed. Inform., vol. 78, pp. 87–101, Feb. 2018, doi: 10.1016/j.jbi.2018.01.004. DOI: https://doi.org/10.1016/j.jbi.2018.01.004

[2] M. B. Gesicho, M. C. Were, and A. Babic, “Data cleaning process for HIV-indicator data extracted from DHIS2 national reporting system: a case study of Kenya,” BMC Med. Inform. Decis. Mak., vol. 20, no. 1, Dec. 2020, doi: 10.1186/s12911-020-01315-7. DOI: https://doi.org/10.1186/s12911-020-01315-7

[3] R. Khare et al., “A longitudinal analysis of data quality in a large pediatric data research network,” Journal of the American Medical Informatics Association, vol. 24, no. 6, Nov. 2017, doi: 10.1093/jamia/ocx033. DOI: https://doi.org/10.1093/jamia/ocx033

[4] E. Albu et al., “Challenges and recommendations for Electronic Health Records data extraction and preparation for dynamic prediction modelling in hospitalized patients -- a practical guide,” Mar. 2025, [Online]. Available: http://arxiv.org/abs/2501.10240 DOI: https://doi.org/10.2196/73987

[5] J. F. Ethier et al., “A unified structural/terminological interoperability framework based on lexEVS: Application to TRANSFoRm,” Journal of the American Medical Informatics Association, vol. 20, no. 5, pp. 986–994, 2013, doi: 10.1136/amiajnl-2012-001312. DOI: https://doi.org/10.1136/amiajnl-2012-001312

[6] Z. Sajjadnia, R. Khayami, and M. R. Moosavi, “Preprocessing Breast Cancer Data to Improve the Data Quality, Diagnosis Procedure, and Medical Care Services,” Cancer Inform., vol. 19, 2020, doi: 10.1177/1176935120917955. DOI: https://doi.org/10.1177/1176935120917955

[7] J. Gaspar, E. Catumbela, B. Marques, and A. Freitas, “A systematic review of outliers detection techniques in medical data: Preliminary study,” in HEALTHINF 2011 - Proceedings of the International Conference on Health Informatics, 2011, pp. 575–582. doi: 10.5220/0003168705750582. DOI: https://doi.org/10.5220/0003168705750582

[8] J. Boone-Heinonen et al., “Not so implausible: impact of longitudinal assessment of implausible anthropometric measures on obesity prevalence and weight change in children and adolescents,” Ann. Epidemiol., vol. 31, pp. 69-74.e5, Mar. 2019, doi: 10.1016/j.annepidem.2019.01.006. DOI: https://doi.org/10.1016/j.annepidem.2019.01.006

[9] Z. Zhao, R. Wang, D. Huang, and Z. Li, “Outlier detection for partially labeled categorical data based on conditional information entropy,” International Journal of Approximate Reasoning, vol. 164, p. 109086, 2024, doi: https://doi.org/10.1016/j.ijar.2023.109086. DOI: https://doi.org/10.1016/j.ijar.2023.109086

[10] Z. Wang, J. R. Talburt, N. Wu, S. Dagtas, and M. N. Zozus, “A Rule-Based Data Quality Assessment System for Electronic Health Record Data,” Appl. Clin. Inform., vol. 11, no. 4, pp. 622–634, Aug. 2020, doi: 10.1055/s-0040-1715567. DOI: https://doi.org/10.1055/s-0040-1715567

[11] L. Zhang et al., “Probabilistic-mismatch anomaly detection: Do one’s medications match with the diagnoses,” in Proceedings - IEEE International Conference on Data Mining, ICDM, Institute of Electrical and Electronics Engineers Inc., Jul. 2016, pp. 659–668. doi: 10.1109/ICDM.2016.12. DOI: https://doi.org/10.1109/ICDM.2016.0077

[12] A. Lighterness, M. Adcock, L. A. Scanlon, and G. Price, “Data Quality–Driven Improvement in Health Care: Systematic Literature Review,” 2024, JMIR Publications Inc. doi: 10.2196/57615. DOI: https://doi.org/10.2196/preprints.57615

[13] A. Jazayeri, O. S. Liang, and C. C. Yang, “Imputation of Missing Data in Electronic Health Records Based on Patients’ Similarities,” J. Healthc. Inform. Res., vol. 4, no. 3, pp. 295–307, Sep. 2020, doi: 10.1007/s41666-020-00073-5. DOI: https://doi.org/10.1007/s41666-020-00073-5

[14] M. Tabassum, S. Mahmood, A. Bukhari, B. Alshemaimri, A. Daud, and F. Khalique, “Anomaly-based threat detection in smart health using machine learning,” BMC Med. Inform. Decis. Mak., vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12911-024-02760-4. DOI: https://doi.org/10.1186/s12911-024-02760-4

[15] L. K. Tee, C. Chee, H. H. Mohamed, and O. S. Lee, “E-clean: A data cleaning framework for patient data,” in Proceedings - 1st International Conference on Informatics and Computational Intelligence, ICI 2011, 2011, pp. 63–68. doi: 10.1109/ICI.2011.21. DOI: https://doi.org/10.1109/ICI.2011.21

[16] B. W. Putra, “Deep Neural Network Classification for ECG Arrhythmia Detection with Optimized Feature Autoencoder,” Journal of Advanced Research in Applied Mechanics, vol. 133, no. 1, pp. 78–86, Sep. 2025, doi: 10.37934/aram.133.1.7886. DOI: https://doi.org/10.37934/aram.133.1.7886

[17] X. Shi, C. Prins, G. Van Pottelbergh, P. Mamouris, B. Vaes, and B. De Moor, “An automated data cleaning method for Electronic Health Records by incorporating clinical knowledge,” BMC Med. Inform. Decis. Mak., vol. 21, no. 1, Dec. 2021, doi: 10.1186/s12911-021-01630-7. DOI: https://doi.org/10.1186/s12911-021-01630-7

[18] D. Antonelli, G. Bruno, and S. Chiusano, “Anomaly detection in medical treatment to discover unusual patient management,” IIE Trans. Healthc. Syst. Eng., vol. 3, no. 2, pp. 69–77, 2013, doi: 10.1080/19488300.2013.787564. DOI: https://doi.org/10.1080/19488300.2013.787564

[19] C. Zhang, X. Xiao, and C. Wu, “Medical fraud and abuse detection system based on machine learning,” Int. J. Environ. Res. Public Health, vol. 17, no. 19, pp. 1–11, Oct. 2020, doi: 10.3390/ijerph17197265. DOI: https://doi.org/10.3390/ijerph17197265

[20] C. Daymont, M. E. Ross, A. R. Localio, A. G. Fiks, R. C. Wasserman, and R. WGrundmeier, “Automated identification of implausible values in growth data from pediatric electronic health records,” Journal of the American Medical Informatics Association, vol. 24, no. 6, pp. 1080–1087, Nov. 2017, doi: 10.1093/jamia/ocx037. DOI: https://doi.org/10.1093/jamia/ocx037

[21] B. Heude et al., “A big-data approach to producing descriptive anthropometric references: a feasibility and validation study of paediatric growth charts,” Lancet Digit. Health, vol. 1, no. 8, pp. e413–e423, Dec. 2019, doi: 10.1016/S2589-7500(19)30149-9. DOI: https://doi.org/10.1016/S2589-7500(19)30149-9

[22] H. Estiri and S. N. Murphy, “Semi-supervised encoding for outlier detection in clinical observation data,” Comput. Methods Programs Biomed., vol. 181, Nov. 2019, doi: 10.1016/j.cmpb.2019.01.002. DOI: https://doi.org/10.1016/j.cmpb.2019.01.002

[23] A. Christy, M. G. Gandhi, and S. Vaithyasubramanian, “Cluster based outlier detection algorithm for healthcare data,” in Procedia Computer Science, Elsevier B.V., 2015, pp. 209–215. doi: 10.1016/j.procs.2015.04.058. DOI: https://doi.org/10.1016/j.procs.2015.04.058

[24] P. Röchner and F. Rothlauf, “Unsupervised anomaly detection of implausible electronic health records: a real-world evaluation in cancer registries,” BMC Med. Res. Methodol., vol. 23, no. 1, Dec. 2023, doi: 10.1186/s12874-023-01946-0. DOI: https://doi.org/10.1186/s12874-023-01946-0

[25] H. T. T. Phan, F. Borca, D. Cable, J. Batchelor, J. H. Davies, and S. Ennis, “Automated data cleaning of paediatric anthropometric data from longitudinal electronic health records: protocol and application to a large patient cohort,” Sci. Rep., vol. 10, no. 1, Dec. 2020, doi: 10.1038/s41598-020-66925-7. DOI: https://doi.org/10.1038/s41598-020-66925-7

[26] Z. Hamid, F. Khalique, S. Mahmood, A. Daud, A. Bukhari, and B. Alshemaimri, “Healthcare insurance fraud detection using data mining,” BMC Med. Inform. Decis. Mak., vol. 24, no. 1, Dec. 2024, doi: 10.1186/s12911-024-02512-4. DOI: https://doi.org/10.1186/s12911-024-02512-4

[27] I. Matloob, S. Khan, R. Rukaiya, H. Alfrahi, and J. Ali Khan, “Healthcare fraud detection using adaptive learning and deep learning techniques,” Evolving Systems, vol. 16, no. 2, Jun. 2025, doi: 10.1007/s12530-025-09698-6. DOI: https://doi.org/10.1007/s12530-025-09698-6

[28] X. Shi, C. Prins, G. Van Pottelbergh, P. Mamouris, B. Vaes, and B. De Moor, “An automated data cleaning method for Electronic Health Records by incorporating clinical knowledge,” BMC Med. Inform. Decis. Mak., vol. 21, no. 1, Dec. 2021, doi: 10.1186/s12911-021-01630-7. DOI: https://doi.org/10.1186/s12911-021-01630-7

[29] B. W. Putra, M. Fachrurrozi, M. R. Sanjaya, A. Muliawati, A. N. S. Mukti, and S. Nurmaini, “Abnormality Heartbeat Classification of ECG Signal Using Deep Neural Network and Autoencoder,” in 2019 International Conference on Informatics, Multimedia, Cyber and Information System (ICIMCIS), IEEE, 2019, pp. 213–218. DOI: https://doi.org/10.1109/ICIMCIS48181.2019.8985206

[30] H. Estiri, J. G. Klann, and S. N. Murphy, “A clustering approach for detecting implausible observation values in electronic health records data,” BMC Med. Inform. Decis. Mak., vol. 19, no. 1, Jul. 2019, doi: 10.1186/s12911-019-0852-6. DOI: https://doi.org/10.1186/s12911-019-0852-6

[31] N. Martin, A. Martinez-Millana, B. Valdivieso, and C. Fernández-Llatas, “Interactive Data Cleaning for Process Mining: A Case Study of an Outpatient Clinic’s Appointment System,” in Business Process Management Workshops, C. Di Francescomarino, R. Dijkman, and U. Zdun, Eds., Cham: Springer International Publishing, 2019, pp. 532–544. DOI: https://doi.org/10.1007/978-3-030-37453-2_43

Downloads

Published

2026-07-31

How to Cite

[1]
B. W. . Putra, M. R. . Sanjaya, and H. Afif, “An Integrated Data Cleaning and Anomaly Detection Pipeline for Improving Clinical Data Quality in Primary Healthcare”, International Journal of Advances in Data and Information Systems, vol. 7, no. 2, pp. 520–533, Jul. 2026, doi: 10.59395/ijadis.v7i2.1485.

Share



Plum Analytics


Similar Articles

11-20 of 154

You may also start an advanced similarity search for this article.