Skip to content
Data Science & information systems International Journal of Advances in Data and Information Systems
Open access E-ISSN 2721-3056 Acceptance rate: 28%

Quantifying Cross-Lingual Sentiment Prediction Difference in Bilingual Pop-Culture Reviews via Dual Monolingual Transformers and Jensen-Shannon Divergence

Authors

  • Azfani Naurotul Jannah Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia image/svg+xml
  • Triando Hamonangan Saragih Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia image/svg+xml
  • Irwan Budiman Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia image/svg+xml
  • Muliadi Muliadi Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia image/svg+xml
  • Radityo Adi Nugroho Faculty of Mathematics and Natutal Science, Department of Computer Science, Lambung Mangkurat University, Kalimantan, Indonesia image/svg+xml

DOI:

https://doi.org/10.59395/ijadis.v7i2.1678

Keywords:

Cross-lingual sentiment analysis, Machine translation, Chinese-RoBERTa, IndoBERT , Jensen–Shannon Divergence

Abstract

Machine translation has been widely used to support cross-lingual sentiment analysis when labeled data in the target language are limited. However, differences in linguistic representation between the source text and its translation may affect the probabilistic outputs of sentiment classification models. This study investigated differences in sentiment predictions between original Mandarin reviews and their Indonesian translations using two independently fine-tuned monolingual transformer models. Chinese-RoBERTa-wwm-ext was applied to the original Mandarin reviews, whereas IndoBERT was used to classify the Indonesian translations. Jensen–Shannon Divergence was employed to compare the sentiment probability distributions generated by the two models. The results showed that Chinese-RoBERTa achieved an accuracy of 74%, whereas IndoBERT achieved 69% on the pseudo-labeled evaluation dataset. Furthermore, 73.81% of the review pairs retained consistent sentiment predictions, while 26.19% exhibited prediction shifts, with Polarity Amplification being the most frequently observed category and most transitions occurring between adjacent sentiment classes. The probability-distribution analysis also revealed substantial differences in prediction confidence for some review pairs, even when the predicted sentiment labels remained identical. These findings demonstrated that comparing probability distributions provided complementary information beyond label-based evaluation for analyzing prediction differences between independently trained monolingual sentiment models on bilingual review pairs. 

392 98

Downloads

Download data is not yet available.

References

[1] Y. Mao, Q. Liu, and Y. Zhang, Sentiment analysis methods, applications, and challenges: A systematic literature review, Apr. 01, 2024, King Saud bin Abdulaziz University. doi: 10.1016/j.jksuci.2024.102048. DOI: https://doi.org/10.1016/j.jksuci.2024.102048

[2] K. L. Tan, C. P. Lee, and K. M. Lim, A Survey of Sentiment Analysis: Approaches, Datasets, and Future Research, Apr. 01, 2023, MDPI. doi: 10.3390/app13074550. DOI: https://doi.org/10.3390/app13074550

[3] Y. Xu, H. Cao, W. Du, and W. Wang, A Survey of Cross-lingual Sentiment Analysis: Methodologies, Models and Evaluations, Sep. 01, 2022, Springer. doi: 10.1007/s41019-022-00187-3. DOI: https://doi.org/10.1007/s41019-022-00187-3

[4] M. Krasitskii, G. Sidorov, O. Kolesnikova, L. Chanona-Hernandez, and A. Gelbukh, Multilingual sentiment analysis of summarized texts: a cross-language study of text shortening effects, PeerJ Comput. Sci., vol. 12, 2026, doi: 10.7717/peerj-cs.3406. DOI: https://doi.org/10.7717/peerj-cs.3406

[5] P. Naveen and P. Trojovsk, Overview and challenges of machine translation for contextually appropriate translations, iScience, vol. 27, no. 10, p. 110878, Oct. 2024, doi: 10.1016/j.isci.2024.110878. DOI: https://doi.org/10.1016/j.isci.2024.110878

[6] P. Pib, J. md, J. Steinberger, and A. Mitera, A comparative study of cross-lingual sentiment analysis, Expert Syst. Appl., vol. 247, p. 123247, Aug. 2024, doi: 10.1016/j.eswa.2024.123247. DOI: https://doi.org/10.1016/j.eswa.2024.123247

[7] M. S. U. Miah, M. M. Kabir, T. Bin Sarwar, M. Safran, S. Alfarhood, and M. F. Mridha, A multimodal approach to cross-lingual sentiment analysis with ensemble of transformer and LLM, Sci. Rep., vol. 14, no. 1, p. 9603, Apr. 2024, doi: 10.1038/s41598-024-60210-7. DOI: https://doi.org/10.1038/s41598-024-60210-7

[8] E. M. Mercha and H. Benbrahim, Machine learning and deep learning for sentiment analysis across languages: A survey, Neurocomputing, vol. 531, pp. 195216, Apr. 2023, doi: 10.1016/j.neucom.2023.02.015. DOI: https://doi.org/10.1016/j.neucom.2023.02.015

[9] A. Velankar, H. Patil, and R. Joshi, Mono vs Multilingual BERT for Hate Speech Detection and Text Classification: A Case Study in Marathi, Apr. 2022, doi: 10.1007/978-3-031-20650-4_10. DOI: https://doi.org/10.1007/978-3-031-20650-4_10

[10] M. Karpinska, N. Raj, K. Thai, Y. Song, A. Gupta, and M. Iyyer, DEMETR: Diagnosing Evaluation Metrics for Translation, Oct. 2022, [Online]. Available: http://arxiv.org/abs/2210.13746 DOI: https://doi.org/10.18653/v1/2022.emnlp-main.649

[11] C. Zhao et al., A Systematic Review of Cross-Lingual Sentiment Analysis: Tasks, Strategies, and Prospects, ACM Comput. Surv., vol. 56, no. 7, Oct. 2024, doi: 10.1145/3645106. DOI: https://doi.org/10.1145/3645106

[12] M. Artetxe, V. Goswami, S. Bhosale, A. Fan, and L. Zettlemoyer, Revisiting Machine Translation for Cross-lingual Classification. [Online]. Available: https://spacy.io/models/xx#xx_sent_ud_sm

[13] S. Al Amer, M. Lee, and P. Smith, Cross-lingual Classification of Crisis-related Tweets Using Machine Translation, in International Conference Recent Advances in Natural Language Processing, RANLP, Incoma Ltd, 2023, pp. 2231. doi: 10.26615/978-954-452-092-2_003. DOI: https://doi.org/10.26615/978-954-452-092-2_003

[14] I. Shode, D. Ifeoluwa Adelani, J. Peng, and A. Feldman, NollySenti: Leveraging Transfer Learning and Machine Translation for Nigerian Movie Sentiment Classification, Short Papers. [Online]. Available: https://github.com/IyanuSh/NollySenti

[15] S. Qian, C. Or, F. do Carmo, and D. Kanojia, Evaluating Machine Translation for Emotion-loaded User Generated Content (TransEval4Emo-UGC). [Online]. Available: https://weibo.com/

[16] Y. Li, T. Duan, and L. Zhu, Public Attitudes and Sentiments toward Common Prosperity in China: A Text Mining Analysis Based on Social Media, Applied Sciences (Switzerland), vol. 14, no. 10, May 2024, doi: 10.3390/app14104295. DOI: https://doi.org/10.3390/app14104295

[17] K. Kuligowska and B. Kowalczuk, Pseudo-labeling with transformers for improving Question Answering systems, in Procedia Computer Science, Elsevier B.V., 2021, pp. 11621169. doi: 10.1016/j.procs.2021.08.119. DOI: https://doi.org/10.1016/j.procs.2021.08.119

[18] J. Zhang et al., Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence, Mar. 2023, [Online]. Available: http://arxiv.org/abs/2209.02970

[19] Y. Cui, W. Che, T. Liu, B. Qin, S. Wang, and G. Hu, Revisiting Pre-Trained Models for Chinese Natural Language Processing, Nov. 2020, doi: 10.18653/v1/2020.findings-emnlp.58. DOI: https://doi.org/10.18653/v1/2020.findings-emnlp.58

[20] Z. Tan et al., Neural machine translation: A review of methods, resources, and tools, Jan. 01, 2020, Elsevier B.V. doi: 10.1016/j.aiopen.2020.11.001. DOI: https://doi.org/10.1016/j.aiopen.2020.11.001

[21] W. Jooste, R. Haque, and A. Way, Philipp Koehn: Neural Machine Translation, Machine Translation, vol. 35, no. 2, pp. 289299, Jun. 2021, doi: 10.1007/s10590-021-09277-x. DOI: https://doi.org/10.1007/s10590-021-09277-x

[22] A. S. Hanin and M. Maryam, Sentiment Analysis of Twitter Towards the Free Lunch Program Using the C4.5 Algorithm, International Journal of Advances in Data and Information Systems, vol. 6, no. 1, pp. 3145, Apr. 2025, doi: 10.59395/ijadis.v6i1.1357. DOI: https://doi.org/10.59395/ijadis.v6i1.1357

[23] V. R. Joseph, Optimal ratio for data splitting, Stat. Anal. Data Min., vol. 15, no. 4, pp. 531538, Aug. 2022, doi: 10.1002/sam.11583. DOI: https://doi.org/10.1002/sam.11583

[24] J. H. Syu, M. Fojcik, R. Cupek, and J. C. W. Lin, HTTPS: Heterogeneous Transfer learning for spliT Prediction System evaluated on healthcare data, Information Fusion, vol. 113, p. 102617, Jan. 2025, doi: 10.1016/J.INFFUS.2024.102617. DOI: https://doi.org/10.1016/j.inffus.2024.102617

[25] Q. H. Nguyen et al., Influence of data splitting on performance of machine learning models in prediction of shear strength of soil, Math. Probl. Eng., vol. 2021, 2021, doi: 10.1155/2021/4832864. DOI: https://doi.org/10.1155/2021/4832864

[26] Y. Liu et al., RoBERTa: A Robustly Optimized BERT Pretraining Approach, Jul. 2019, [Online]. Available: http://arxiv.org/abs/1907.11692

[27] F. Qiu, H. Xiao, K. Yu, and Q. Qiu, MP-CNER: A Course Named Entity Recognition Model Based on Multi-Dimensional Position Features, Electronics (Switzerland), vol. 15, no. 11, Jun. 2026, doi: 10.3390/electronics15112396. DOI: https://doi.org/10.3390/electronics15112396

[28] I. Loshchilov and F. Hutter, DECOUPLED WEIGHT DECAY REGULARIZATION. [Online]. Available: https://github.com/loshchil/AdamW-and-SGDW

[29] R. Anadra, H. Wijayanto, and K. Sadik, Sentiment Analysis of Tokopedia Customer Reviews Using BiLSTM and IndoBERT with Comparative Analysis of Preprocessing and Labeling Methods, International Journal of Advances in Data and Information Systems, vol. 6, no. 3, pp. 773788, Dec. 2025, doi: 10.59395/ijadis.v6i3.1458. DOI: https://doi.org/10.59395/ijadis.v6i3.1458

[30] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP, Online. [Online]. Available: https://huggingface.co/

[31] D. Jurafsky and J. H. Martin, Speech and Language Processing An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models Third Edition draft.

[32] M. Owusu-Adjei, J. Ben Hayfron-Acquah, T. Frimpong, and G. Abdul-Salaam, A systematic review of prediction accuracy as an evaluation measure for determining machine learning model performance in healthcare systems, Jun. 04, 2023. doi: 10.1101/2023.06.01.23290837. DOI: https://doi.org/10.1101/2023.06.01.23290837

[33] D. M. W. Powers and Ailab, EVALUATION: FROM PRECISION, RECALL AND F-MEASURE TO ROC, INFORMEDNESS, MARKEDNESS & CORRELATION.

[34] D. J. Hand, P. Christen, and N. Kirielle, F*: an interpretable transformation of the F-measure, Mach. Learn., vol. 110, no. 3, pp. 451456, Mar. 2021, doi: 10.1007/s10994-021-05964-1. DOI: https://doi.org/10.1007/s10994-021-05964-1

[35] B. Shade and E. G. Altmann, Quantifying the Dissimilarity of Texts, Information (Switzerland), vol. 14, no. 5, May 2023, doi: 10.3390/info14050271. DOI: https://doi.org/10.3390/info14050271

[36] X. Zhu, S. Gardiner, T. Roldn, and D. Rossouw, The Model Arena for Cross-lingual Sentiment Analysis: A Comparative Study in the Era of Large Language Models, 2024. [Online]. Available: https://platform.openai.com/docs/models/ DOI: https://doi.org/10.18653/v1/2024.wassa-1.12

[37] J.-Q. Wang, CLAS-Net: A study on cross-lingual intelligent sentiment analysis model fusing semantic alignment, PLoS One, vol. 21, no. 2, p. e0342342, Feb. 2026, doi: 10.1371/journal.pone.0342342. DOI: https://doi.org/10.1371/journal.pone.0342342

Downloads

Published

2026-08-30

How to Cite

[1]
A. N. . Jannah, T. H. Saragih, I. Budiman, M. Muliadi, and R. A. Nugroho, “Quantifying Cross-Lingual Sentiment Prediction Difference in Bilingual Pop-Culture Reviews via Dual Monolingual Transformers and Jensen-Shannon Divergence”, International Journal of Advances in Data and Information Systems, vol. 7, no. 2, pp. 960–977, Aug. 2026, doi: 10.59395/ijadis.v7i2.1678.

Share



Plum Analytics


Similar Articles

21-30 of 126

You may also start an advanced similarity search for this article.