MACHINE LEARNING / NLP
Bengali News Classification
BSc thesis. I scraped Bangladeshi news sites with Scrapy, built a dataset of 504,266 labelled articles in 7 categories, and compared Naive Bayes, SVM, CNN, LSTM and CNN+LSTM models. The best deep models reached 93.3% accuracy on 50,062 held-out articles.
- 504,266 articles
- 7 categories
- 93.3% test accuracy
- Context
- BSc thesis, Daffodil International University
- Period
- 2020 – 2021
Stack
- Python
- Scrapy
- pandas
- scikit-learn
- Keras
- TensorFlow
- Google Colab
Dataset
No large labelled Bengali news corpus was available, so I wrote Scrapy spiders for Bangladeshi news sites, including Jago News, Bangladesh Pratidin and Ittefaq. I used each site's own section as the label.
- articles collected
- 504,266
- categories
- 7
- sports, international, national, all-Bangladesh, politics, entertainment, economics-business
- after cleaning
- 500,620
- held-out test articles
- 50,062
- 10% split
Counts from the notebook outputs in the repository.
Preprocessing
- Removed empty records and articles shorter than 490 or longer than 5,000 characters, which were mostly stubs or merged pages.
- Tokenised Bengali text with a Keras tokenizer and padded sequences to 250 tokens for the CNN and 500 tokens for the LSTM.
- Used TF-IDF and count features for the classical baselines, and one-hot labels for the neural models.
Model experimentation
| Model | Features | Test accuracy |
|---|---|---|
| Naive Bayes | Bag of words | 85% |
| SVM (linear kernel) | TF-IDF | 90.9% |
| CNN | Embeddings, 250 tokens | 93.3% |
| LSTM | Embeddings, 500 tokens | 93.3% |
| CNN + LSTM | Embeddings, 250 tokens (6 classes) | 92.9% |
Training accuracy for the neural models reached 97–98%, while validation accuracy levelled off around 93–94% after two or three epochs. I report test accuracy, and the gap between the two told me where to stop training.
Result
Both the CNN and the LSTM reached 93.3% accuracy with macro F1 of 0.93 across the seven classes. The CNN trained about twice as fast per epoch, which made it the practical choice.
What I learned
- Collecting and cleaning the data took more work than modelling, and it set the upper limit on accuracy.
- Strong classical baselines (SVM at 90.9%) are important context for deep-learning results.
- Reporting held-out metrics rather than training accuracy is the difference between a demo and an evaluation.