Skip to content
All projects

MACHINE LEARNING / NLP

Bengali News Classification

BSc thesis. I scraped Bangladeshi news sites with Scrapy, built a dataset of 504,266 labelled articles in 7 categories, and compared Naive Bayes, SVM, CNN, LSTM and CNN+LSTM models. The best deep models reached 93.3% accuracy on 50,062 held-out articles.

  • 504,266 articles
  • 7 categories
  • 93.3% test accuracy
Context
BSc thesis, Daffodil International University
Period
2020 – 2021

Stack

  • Python
  • Scrapy
  • pandas
  • scikit-learn
  • Keras
  • TensorFlow
  • Google Colab

Dataset

No large labelled Bengali news corpus was available, so I wrote Scrapy spiders for Bangladeshi news sites, including Jago News, Bangladesh Pratidin and Ittefaq. I used each site's own section as the label.

articles collected
504,266
categories
7
sports, international, national, all-Bangladesh, politics, entertainment, economics-business
after cleaning
500,620
held-out test articles
50,062
10% split

Counts from the notebook outputs in the repository.

Preprocessing

  • Removed empty records and articles shorter than 490 or longer than 5,000 characters, which were mostly stubs or merged pages.
  • Tokenised Bengali text with a Keras tokenizer and padded sequences to 250 tokens for the CNN and 500 tokens for the LSTM.
  • Used TF-IDF and count features for the classical baselines, and one-hot labels for the neural models.

Model experimentation

Accuracy on the held-out test set
ModelFeaturesTest accuracy
Naive BayesBag of words85%
SVM (linear kernel)TF-IDF90.9%
CNNEmbeddings, 250 tokens93.3%
LSTMEmbeddings, 500 tokens93.3%
CNN + LSTMEmbeddings, 250 tokens (6 classes)92.9%

Training accuracy for the neural models reached 97–98%, while validation accuracy levelled off around 93–94% after two or three epochs. I report test accuracy, and the gap between the two told me where to stop training.

Result

Both the CNN and the LSTM reached 93.3% accuracy with macro F1 of 0.93 across the seven classes. The CNN trained about twice as fast per epoch, which made it the practical choice.

What I learned

  • Collecting and cleaning the data took more work than modelling, and it set the upper limit on accuracy.
  • Strong classical baselines (SVM at 90.9%) are important context for deep-learning results.
  • Reporting held-out metrics rather than training accuracy is the difference between a demo and an evaluation.

Source