NAÏVE BAYES SENTIMENT ACCURACY WITH CHI-SQUARE AND INFORMATION GAIN

Authors

  • Muhamad Fahrurozi STMIK IKMI CIREBON
  • Dian Ade Kurnia STMIK IKMI Cirebon
  • Yudhistira Arie Wijaya STMIK IKMI Cirebon
  • Puji Pramudya Marta STMIK IKMI Cirebon
  • Khaerul Anam STMIK IKMI Cirebon

DOI:

https://doi.org/10.35457/1nmksc45

Keywords:

Chi-Square, Information Gain, sentiment classification, Naive Bayes, feature selection

Abstract

This study aims to evaluate the effectiveness of the Chi-Square and Information Gain feature selection techniques in improving the accuracy of sentiment classification on application reviews using the Naive Bayes algorithm. The main issue addressed in this research is the high dimensionality of features in Indonesian-language text data, which may potentially affect model performance. To address this, the study applies a text preprocessing pipeline consisting of sentiment labeling, normalization, and feature extraction using TF-IDF with an initial 5000 features. The dataset contains 225,043 Gojek application reviews, which were reduced to 220,860 valid entries after cleaning, and subsequently divided into training and testing sets using a stratified split. Experimental results show that Chi-Square achieved the highest accuracy of 0.877864 with 3000 features, while Information Gain reached an accuracy of 0.877751 with 4000 features. Both values are slightly lower than the model without feature selection, which achieved an accuracy of 0.878045. These findings indicate that feature reduction does not improve the performance of Naive Bayes, as the algorithm performs more effectively when retaining a broader distribution of words and more complete contextual representations. In conclusion, feature selection for probabilistic models should be applied cautiously, especially on Indonesian-language review data that are informal and exhibit high variation in expression.

Downloads

Download data is not yet available.

References

[1] H. Duong, “User Perception Analysis in Digital Application Ecosystems,” J. Digit. Interact. Stud., vol. 12, no. 1, pp. 45–58, 2024, doi: 10.1016/j.jdis.2024.01.004.

[2] L. Zhang, “Advances in Sentiment Analysis for Large-Scale Text Data,” Int. J. Comput. Linguist., vol. 9, no. 2, pp. 77–93, 2020, doi: 10.1186/s40104-020-00452-7.

[3] A. McCallum, “A Comparison of Event Models for Naive Bayes Text Classification,” in AAAI-98 Workshop on Learning for Text Categorization, 1998, pp. 41–48.

[4] Y. Liu, M. Chen, and J. Xu, “Improving text classification using feature selection with Chi-Square and Information Gain,” Appl. Intell., vol. 53, pp. 14212–14228, 2023, doi: 10.1007/s10489-023-04576-0.

[5] W. Tan, “Evaluating Information Gain for Modern Text Mining Pipelines,” Data Sci. Rev., vol. 14, no. 3, pp. 201–215, 2023, doi: 10.1145/3591123.

[6] Q. Wen, “Improving Sentiment Classification Accuracy with Feature Filtering,” J. Intell. Syst., vol. 16, no. 2, pp. 89–102, 2024, doi: 10.1515/jisys-2023-0047.

[7] C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge: Cambridge University Press, 2008.

[8] R. Churchill and L. Singh, “textPrep: A Text Preprocessing Toolkit for Topic Modeling on Social Media Data ,” 2021. doi: 10.5220/0010559000600070 .

[9] M. Siino, I. Tinnirello, and M. La Cascia, “Is text preprocessing still worth the time? A comparative survey on the influence of popular preprocessing methods on Transformers and traditional classifiers ,” 2024. doi: 10.1016/j.is.2023.102342 .

[10] I. M. Mubaroq and E. B. Setiawan, “The Effect of Information Gain Feature Selection for Hoax Identification in Twitter Using Classification Method Support Vector Machine ,” 2020. doi: 10.34818/INDOJC.2020.5.2.499 .

[11] D. C. Sami’un, A. Sugiharto, and F. Jie, “Chi-Square feature selection for improving sentiment analysis of news data privacy treats ,” 2024. doi: (tidak ada DOI pada versi PDF, tetapi artikel tersedia di JATIT) .

[12] M. T. F. Maulana, B. I. Nugroho, and E. U. S. Utami, “Analisis Sentimen Ulasan Aplikasi TikTok di Google Playstore Menggunakan Algoritma Naïve Bayes,” RIGGS J. Artif. Intell. Digit. Bus., vol. 4, no. 3, 2025, doi: 10.31004/riggs.v4i3.2962.

[13] S. Ulyaa, A. Ridwan, and W. C. Wahyudi, “Text Mining Sentimen Analisis Pengguna Aplikasi Marketplace Tokopedia Berdasarkan Rating dan Komentar pada Google Play Store,” J. Bisnis Digit. dan Sist. Inf., pp. 33–40, 2022, [Online]. Available: https://ejr.umku.ac.id/index.php/BIDISFO/article/download/1799/1071

[14] R. Zulfiqri, B. N. Sari, and T. N. Padilah, “Analisis Sentimen Ulasan Pengguna Aplikasi Media Sosial Instagram pada Google Play Store Menggunakan Naïve Bayes Classifier,” J. Inform. dan Tek. Elektro Terap., vol. 12, no. 3, 2024, doi: 10.23960/jitet.v12i3.4995.

[15] W. Zhang, X. Li, Y. Deng, L. Bing, and W.-K. Lam, “A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges ,” 2022. doi: 10.1109/TKDE.2022.3230975 .

[16] R. Helma, “Sentiment Analysis of Indonesian Mobile Application Reviews,” Indones. J. Data Sci., vol. 4, no. 2, pp. 101–118, 2025, doi: 10.51595/ijds.2025.4.2.101.

[17] M. Zakaria, “Sentiment Analysis of User Feedback in Digital Applications,” J. Inf. Syst. Technol., vol. 12, no. 1, pp. 55–68, 2024, doi: 10.5829/jist.2024.12.1.55.

[18] S. Khoerunnisa, D. F. Shiddieq, and D. Nurhayati, “Penerapan Algoritma Naive Bayes dengan Teknik TF-IDF dan Cross Validation untuk Analisis Sentimen Terhadap Starlink,” MALCOM Indones. J. Mach. Learn. Comput. Sci., vol. 5, no. 2, 2025, doi: 10.57152/malcom.v5i2.1852.

[19] A. D. M. Putri, N. Sulistianingsih, and R. Rismayati, “Pengaruh Teknik Representasi Teks Bag-of-Words dan TF-IDF terhadap Akurasi Klasifikasi Sentimen Teks Multi-Domain,” JTIM J. Teknol. Inf. dan Multimed., vol. 7, no. 4, 2025, doi: 10.35746/jtim.v7i4.756.

[20] W. Tan, “Evaluating N-gram and TF-IDF Combinations for Text Classification,” Data Sci. Rev., vol. 14, no. 3, pp. 201–215, 2023, doi: 10.1145/3591123.

[21] Q. Wen, “Improving Feature Selection for Sentiment Classification Using Information Gain,” J. Intell. Syst., vol. 16, no. 2, pp. 89–102, 2024, doi: 10.1515/jisys-2023-0047.

[22] S. A. Anjani and A. Fauzan, “Implementasi n-Gram dalam Analisis Sentimen Masyarakat DIY terhadap PSBB Jawa-Bali,” Statistika, vol. 21, no. 2, 2021, doi: 10.29313/statistika.v21i2.294.

[23] A. Solikhatun and E. Sugiharti, “Application of the Naïve Bayes Classifier Algorithm using N-Gram and Information Gain to Improve the Accuracy of Restaurant Review Sentiment Analysis,” J. Adv. Inf. Syst. Technol., vol. 2, no. 2, pp. 11–20, 2020, doi: 10.15294/jaist.v2i2.44303.

[24] N. Fadhilah, Y. Prasetyo, and A. Wibowo, “Peningkatan Akurasi Klasifikasi Teks melalui Seleksi Fitur Information Gain dan Chi-Square,” J. Sist. Inf., vol. 19, no. 2, pp. 101–110, 2023, doi: 10.24089/jsi.v19i2.4567.

[25] A. Ramadhan and R. A. Asmara, “Penerapan Metode Seleksi Fitur Chi-Square pada Klasifikasi Sentimen Menggunakan Naïve Bayes,” J. Teknol. Inf. dan Komput., vol. 10, no. 1, pp. 55–63, 2022, doi: 10.33365/jtik.v10i1.2431.

[26] Y. Liu, “Chi-Square Based Feature Selection for Text Classification,” J. Mach. Learn. Appl., vol. 5, no. 4, pp. 120–137, 2023, doi: 10.32604/jmla.2023.024111.

[27] J. R. Quinlan, C4.5: Programs for Machine Learning. San Francisco: Morgan Kaufmann, 1993.

[28] H. Allam, M. Yaghoobi, and H. X. Nguyen, “Text Classification: How Machine Learning Is Used to Categorize Text Data ,” 2025. doi: 10.3390/info16020130 .

[29] K. Taha, “A Comprehensive Survey of Modern Text Classification Techniques ,” 2024.

[30] D. Jurafsky and J. H. Martin, Speech and Language Processing. Upper Saddle River, NJ: Prentice Hall, 2023. [Online]. Available: https://web.stanford.edu/~jurafsky/slp3/

[31] D. Nurhayati, “Sentiment Analysis of Indonesian App Reviews using Naive Bayes and SVM,” J. Inform. Indones., vol. 5, no. 2, pp. 75–86, 2019, doi: 10.25008/jii.2019.5.2.75.

[32] R. Singh, “Enhanced Naive Bayes with Feature Selection for Sentiment Analysis,” Int. J. Data Sci., vol. 7, no. 1, pp. 44–58, 2020, doi: 10.5120/ijds.2020.71.44.

Downloads

Published

2026-05-31

Issue

Section

Articles

How to Cite

[1]
“NAÏVE BAYES SENTIMENT ACCURACY WITH CHI-SQUARE AND INFORMATION GAIN”, antivirus, vol. 20, no. 1, pp. 98–109, May 2026, doi: 10.35457/1nmksc45.

Similar Articles

1-10 of 99

You may also start an advanced similarity search for this article.