Skip to content
Data Science & information systems International Journal of Advances in Data and Information Systems
Open access E-ISSN 2721-3056 Acceptance rate: 28%

Sentiment Analysis of Tokopedia Customer Reviews Using BiLSTM and IndoBERT with Comparative Analysis of Preprocessing and Labeling Methods

Authors

  • Rahmi Anadra IPB University
  • Hari Wijayanto IPB University
  • Kusman Sadik IPB University

DOI:

https://doi.org/10.59395/ijadis.v6i3.1458

Keywords:

Sentiment analysis, BiLSTM, IndoBERT, Text preprocessing, Data labeling

Abstract

This study addresses key challenges in Indonesian sentiment analysis related to preprocessing, labeling strategies, and class imbalance. It compares the performance of BiLSTM and IndoBERT using user reviews collected from Tokopedia. The dataset was manually and automatically labeled, then processed under three preprocessing schemes. Both models were trained with tuned hyperparameters and imbalance-handling techniques and evaluated through twenty rounds of stratified five-fold cross-validation. Performance was assessed using balanced accuracy and F1-score. IndoBERT achieved the highest results, with balanced accuracy up to 0.85 and F1-scores up to 0.83, while BiLSTM reached balanced accuracy up to 0.78 and F1-scores up to 0.76. Applying class weight and focal loss improved model performance by approximately 2% to 11% over the baseline. BiLSTM demonstrated greater training efficiency, requiring only 1 to 2.5 minutes per epoch, compared with IndoBERT’s 2.6 to 3.6 minutes. Although manual labeling remained superior in capturing contextual nuance and emotional cues, GPT-based labeling showed strong agreement with the human annotations. A four-way ANOVA revealed that all main factors and several interactions significantly influenced classification outcomes. Overall, BiLSTM provides faster training efficiency, whereas IndoBERT delivers higher predictive accuracy.

2354 1576

Downloads

Download data is not yet available.

References

[1] E. F. Santika, ECDB: Proyeksi Pertumbuhan e-Commerce Indonesia Tertinggi Sedunia pada 2024, Databoks, 2024. https://databoks.katadata.co.id/datapublish/2024/04/29/ecdb-proyeksi-pertumbuhan-e-commerce-indonesia-tertinggi-sedunia-pada-2024 (accessed May 17, 2024).

[2] T. Chen, P. Samaranayake, X. Y. Cen, M. Qi, and Y. C. Lan, The Impact of Online Reviews on Consumers Purchasing Decisions: Evidence From an Eye-Tracking Study, Front. Psychol., vol. 13, no. June, 2022, doi: 10.3389/fpsyg.2022.865702. DOI: https://doi.org/10.3389/fpsyg.2022.865702

[3] Y. Ahn and J. Lee, The Impact of Online Reviews on Consumers Purchase Intentions: Examining the Social Influence of Online Reviews, Group Similarity, and Self-Construal, J. Theor. Appl. Electron. Commer. Res. , vol. 19, no. 2, pp. 10601078, 2024, doi: 10.3390/jtaer19020055. DOI: https://doi.org/10.3390/jtaer19020055

[4] C. H. Lin and U. Nuha, Sentiment analysis of Indonesian datasets based on a hybrid deep-learning strategy, J. Big Data, vol. 10, no. 1, 2023, doi: 10.1186/s40537-023-00782-9. DOI: https://doi.org/10.1186/s40537-023-00782-9

[5] L. Wang, G. Che, J. Hu, and L. Chen, Online Review Helpfulness and Information Overload: The Roles of Text, Image, and Video Elements, J. Theor. Appl. Electron. Commer. Res. , vol. 19, no. 2, pp. 12431266, 2024, doi: 10.3390/jtaer19020064. DOI: https://doi.org/10.3390/jtaer19020064

[6] K. L. Tan, C. P. Lee, and K. M. Lim, A Survey of Sentiment Analysis: Approaches, Datasets, and Future Research, Appl. Sci., vol. 13, no. 7, 2023, doi: 10.3390/app13074550. DOI: https://doi.org/10.3390/app13074550

[7] R. N. Amalia, K. Sadik, and K. A. Notodiputro, a Preliminary Study of Sentiment Analysis on Covid-19 News: Lesson Learned From Data Acquisition, Pre-Processing, and Descriptive Analytics, Barekeng, vol. 17, no. 4, pp. 19011914, 2023, doi: 10.30598/barekengvol17iss4pp1901-1914. DOI: https://doi.org/10.30598/barekengvol17iss4pp1901-1914

[8] M. Wu, Z. Wen, P. Xiao, and W. Li, Performance Comparison of CNN and BiLSTM with FastText Word Embeddings for Chinese Weibo Sentiment Analysis, vol. 13, no. 7, pp. 100106, 2025.

[9] C. Suhaeni, H. Wijayanto, and A. Kurnia, Sentiment Classification on the 2024 Indonesian Presidential Candidate Dataset Using Deep Learning Approaches, Indones. J. Stat. Its Appl., vol. 8, no. 2, pp. 8394, 2024, doi: 10.29244/ijsa.v8i2p83-94. DOI: https://doi.org/10.29244/ijsa.v8i2p83-94

[10] N. K. Nissa and E. Yulianti, Multi-label text classification of Indonesian customer reviews using bidirectional encoder representations from transformers language model, Int. J. Electr. Comput. Eng., vol. 13, no. 5, pp. 56415652, 2023, doi: 10.11591/ijece.v13i5.pp5641-5652. DOI: https://doi.org/10.11591/ijece.v13i5.pp5641-5652

[11] F. Baharuddin and M. F. Naufal, Fine-Tuning IndoBERT for Indonesian Exam Question Classification Based on Blooms Taxonomy, J. Inf. Syst. Eng. Bus. Intell., vol. 9, no. 2, pp. 253263, 2023, doi: 10.20473/jisebi.9.2.253-263. DOI: https://doi.org/10.20473/jisebi.9.2.253-263

[12] M. Siino, I. Tinnirello, and M. La Cascia, Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers, Inf. Syst., vol. 121, no. March 2023, p. 102342, 2024, doi: 10.1016/j.is.2023.102342. DOI: https://doi.org/10.1016/j.is.2023.102342

[13] M. A. Palomino and F. Aider, Evaluating the Effectiveness of Text Pre-Processing in Sentiment Analysis, Appl. Sci., vol. 12, no. 17, 2022, doi: 10.3390/app12178765. DOI: https://doi.org/10.3390/app12178765

[14] S. Biswas, K. Young, and J. Griffith, A Comparison of Automatic Labelling Approaches for Sentiment Analysis, no. Data, pp. 312319, 2022, doi: 10.5220/0011265900003269. DOI: https://doi.org/10.5220/0011265900003269

[15] F. Gilardi, M. Alizadeh, and M. Kubli, ChatGPT outperforms crowd workers for text-annotation tasks, Proc. Natl. Acad. Sci. U. S. A., vol. 120, no. 30, pp. 13, 2023, doi: 10.1073/pnas.2305016120. DOI: https://doi.org/10.1073/pnas.2305016120

[16] A. N. Maaly, D. Pramesti, A. D. Fathurahman, and H. Fakhrurroja, Exploring Sentiment Analysis for the Indonesian Presidential Election Through Online Reviews Using Multi-Label Classification with a Deep Learning Algorithm, Inf., vol. 15, no. 11, pp. 133, 2024, doi: 10.3390/info15110705. DOI: https://doi.org/10.3390/info15110705

[17] Y. Asri, D. Kuswardani, W. N. Suliyanti, Y. O. Manullang, and A. R. Ansyari, Sentiment analysis based on Indonesian language lexicon and IndoBERT on user reviews PLN mobile application, Indones. J. Electr. Eng. Comput. Sci., vol. 38, no. 1, p. 677, 2025, doi: 10.11591/ijeecs.v38.i1.pp677-688. DOI: https://doi.org/10.11591/ijeecs.v38.i1.pp677-688

[18] T. Cai and X. Zhang, Multi Channel BLTCN BLSTM Self Attention, Sensors, pp. 215, 2023.

[19] M. F. Juna and M. Hayaty, The observed preprocessing strategies for doing automatic text summarizing, Comput. Sci. Inf. Technol., vol. 4, no. 2, pp. 119126, 2023, doi: 10.11591/csit.v4i2.pp119-126. DOI: https://doi.org/10.11591/csit.v4i2.pp119-126

[20] M. A. Rosid, A. S. Fitrani, I. R. I. Astutik, N. I. Mulloh, and H. A. Gozali, Improving Text Preprocessing for Student Complaint Document Classification Using Sastrawi, IOP Conf. Ser. Mater. Sci. Eng., vol. 874, no. 1, 2020, doi: 10.1088/1757-899X/874/1/012017. DOI: https://doi.org/10.1088/1757-899X/874/1/012017

[21] L. Zhu and D. Luo, A Novel Efficient and Effective Preprocessing Algorithm for Text Classification, J. Comput. Commun., vol. 11, no. 03, pp. 114, 2023, doi: 10.4236/jcc.2023.113001. DOI: https://doi.org/10.4236/jcc.2023.113001

[22] A. Conneau et al., Unsupervised Cross-lingual Representation Learning at Scale, 2020. DOI: https://doi.org/10.21437/Interspeech.2021-329

[23] G. Xu, Z. Chen, and Z. Zhang, Aspect category sentiment analysis based on pre-trained BiLSTM and syntax-aware graph attention network, Sci. Rep., vol. 15, no. 1, pp. 115, 2025, doi: 10.1038/s41598-025-86009-8. DOI: https://doi.org/10.1038/s41598-025-86009-8

[24] C. Liu, Long short-term memory (LSTM)-based news classification model, PLoS One, vol. 19, no. 5 May, pp. 123, 2024, doi: 10.1371/journal.pone.0301835. DOI: https://doi.org/10.1371/journal.pone.0301835

[25] F. Alkomah and X. Ma, A Literature Review of Textual Hate Speech Detection Methods and Datasets, Inf., vol. 13, no. 6, pp. 122, 2022, doi: 10.3390/info13060273. DOI: https://doi.org/10.3390/info13060273

[26] K. Chatsiou, Text Classification of Manifestos and COVID-19 Press Briefings using BERT and Convolutional Neural Networks, pp. 112, 2020, [Online]. Available: http://arxiv.org/abs/2010.10267

[27] A. Vaswani et al., Attention is all you need, Adv. Neural Inf. Process. Syst., vol. 2017-Decem, no. Nips, pp. 59996009, 2017.

[28] M. K. Shaik Vadla, M. A. Suresh, and V. K. Viswanathan, Enhancing Product Design through AI-Driven Sentiment Analysis of Amazon Reviews Using BERT, Algorithms, vol. 17, no. 2, p. 59, 2024, doi: 10.3390/a17020059. DOI: https://doi.org/10.3390/a17020059

[29] Anugerah Simanjuntak et al., Research and Analysis of IndoBERT Hyperparameter Tuning in Fake News Detection, J. Nas. Tek. Elektro dan Teknol. Inf., vol. 13, no. 1, pp. 6067, 2024, doi: 10.22146/jnteti.v13i1.8532. DOI: https://doi.org/10.22146/jnteti.v13i1.8532

[30] W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, A survey on imbalanced learning: latest research, applications and future directions, vol. 57, no. 6. 2024. doi: 10.1007/s10462-024-10759-6. DOI: https://doi.org/10.1007/s10462-024-10759-6

[31] M. N. Razali, N. Arbaiy, P. C. Lin, and S. Ismail, Optimizing Multiclass Classification Using Convolutional Neural Networks with Class Weights and Early Stopping for Imbalanced Datasets, Electron., vol. 14, no. 4, pp. 114, 2025, doi: 10.3390/electronics14040705. DOI: https://doi.org/10.3390/electronics14040705

[32] T. Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollar, Focal Loss for Dense Object Detection, IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 2, pp. 318327, 2020, doi: 10.1109/TPAMI.2018.2858826. DOI: https://doi.org/10.1109/TPAMI.2018.2858826

[33] Y. N. Kunang, S. Nurmaini, D. Stiawan, and B. Y. Suprapto, Deep learning with focal loss approach for attacks classification, Telkomnika (Telecommunication Comput. Electron. Control., vol. 19, no. 4, pp. 14071418, 2021, doi: 10.12928/TELKOMNIKA.v19i4.18772. DOI: https://doi.org/10.12928/telkomnika.v19i4.18772

[34] I. Markoulidakis, I. Rallis, I. Georgoulas, G. Kopsiaftis, A. Doulamis, and N. Doulamis, Multiclass Confusion Matrix Reduction Method and Its Application on Net Promoter Score Classification Problem, Technologies, vol. 9, no. 4, 2021, doi: 10.3390/technologies9040081. DOI: https://doi.org/10.3390/technologies9040081

[35] S. Farhadpour, T. A. Warner, and A. E. Maxwell, Selecting and Interpreting Multiclass Loss and Accuracy Assessment Metrics for Classifications with Class Imbalance: Guidance and Best Practices, Remote Sens., vol. 16, no. 3, pp. 122, 2024, doi: 10.3390/rs16030533. DOI: https://doi.org/10.3390/rs16030533

[36] K. Nand Kumar, Mikailalsys Journal of, Mikailalsys J. Math. Stat., vol. 2, no. 3, pp. 102113, 2024. DOI: https://doi.org/10.58578/mjms.v2i3.3449

[37] A. S. Rizkia, Wufron, and F. F. Roji, Sentiment Analysis of Coretax: A Comparison of Manual , Transformers- Based , and Lexicon-Based Data Labeling on IndoBERT Performance Analisis Sentimen Coretax: Perbandingan Pelabelan Data Manual , Indones. J. Mach. Learn. Comput. Sci., vol. 5, no. July, pp. 10371048, 2025. DOI: https://doi.org/10.57152/malcom.v5i3.2151

Downloads

Published

2025-12-01

How to Cite

[1]
R. Anadra, H. Wijayanto, and K. Sadik, “Sentiment Analysis of Tokopedia Customer Reviews Using BiLSTM and IndoBERT with Comparative Analysis of Preprocessing and Labeling Methods”, International Journal of Advances in Data and Information Systems, vol. 6, no. 3, pp. 773–788, Dec. 2025, doi: 10.59395/ijadis.v6i3.1458.

Share



Plum Analytics


Similar Articles

1-10 of 172

You may also start an advanced similarity search for this article.