Evaluation of Classical Machine Learning Methods and TextCNN in Indonesian News Classification
DOI:
https://doi.org/10.24235/k0d8my33Keywords:
News Classification, Indonesian Language, Logistic Regression, TextCNN, Small-scale CorpusAbstract
The massive growth of digital news demands reliable automated classification systems. However, systematic evaluations of optimal methods for the Indonesian language, particularly under data-constrained environments, remain scarce. This study evaluates and compares conventional Machine Learning and Deep Learning approaches in classifying multiclass Indonesian news articles. Utilizing a dataset of 1,199 articles scraped from Detik.com, this research applies an automated labeling scheme based on URL subdomain structures. Three methods are comparatively tested: TF-IDF-based Multinomial Naïve Bayes (MNB) and Logistic Regression (LR), against a Deep Learning TextCNN architecture, using a combined title and three-paragraph feature representation. Experimental results demonstrate that Logistic Regression achieves the highest performance (86.67% accuracy, 78.84% F1-Macro), significantly outperforming TextCNN (67.08% accuracy) due to data scarcity constraints. Error analysis reveals that model misclassifications are concentrated within semantically close categories, such as misclassifying Finance as News due to lexical overlap in government policy discussions. These findings underscore that for small-scale Indonesian corpora, classical TF-IDF methods offer superior classification efficacy and computational efficiency compared to complex Deep Learning architectures.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Gina Nufus, Muhammad Rafly Saputra, Dika Hasan Nugroho, Muhammad Syifaulqulub Hakim (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.



