Research Article

An Ensemble-Driven Hybrid Framework for Fine-Grained Bengali Cyberbullying Detection Across Classical and Deep Learning Paradigms

by  Syed Jakir Ahmed
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 127
Published: July 2026
Authors: Syed Jakir Ahmed
10.5120/ijca3ce192fe8d22
PDF

Syed Jakir Ahmed . An Ensemble-Driven Hybrid Framework for Fine-Grained Bengali Cyberbullying Detection Across Classical and Deep Learning Paradigms. International Journal of Computer Applications. 187, 127 (July 2026), 1-10. DOI=10.5120/ijca3ce192fe8d22

                        @article{ 10.5120/ijca3ce192fe8d22,
                        author  = { Syed Jakir Ahmed },
                        title   = { An Ensemble-Driven Hybrid Framework for Fine-Grained Bengali Cyberbullying Detection Across Classical and Deep Learning Paradigms },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 127 },
                        pages   = { 1-10 },
                        doi     = { 10.5120/ijca3ce192fe8d22 },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Syed Jakir Ahmed
                        %T An Ensemble-Driven Hybrid Framework for Fine-Grained Bengali Cyberbullying Detection Across Classical and Deep Learning Paradigms%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 127
                        %P 1-10
                        %R 10.5120/ijca3ce192fe8d22
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

Cyberbullying on Bengali social media platforms has grown into a serious societal concern, yet automated detection remains challenging due to the morphological richness and low-resource nature of the Bengali language. Existing studies have largely relied on binary classification, single-platform corpora, or computationally expensive transformer architectures, leaving a clear gap in robust multiclass detection frameworks that balance accuracy with practical deployability. There is a pressing need for a systematic approach that exploits the complementary strengths of classical and neural models within a unified pipeline. This study proposes a hybrid ensemble framework combining Random Forest, Multi-Layer Perceptron, Convolutional Neural Network, and Recurrent Neural Network as base learners, integrated through soft voting, hard voting, and a CNN-based stacking meta-learner. Experiments were conducted on a publicly available Bengali cyberbullying dataset of 6,005 samples, augmented to 12,500 balanced instances across five abuse categories— Political, Troll, Sexual, Threat, and Neutral. A comparative evaluation of fourteen models demonstrates that the standalone CNN achieves 92.6% accuracy with a ROC AUC of 0.989, outperforming prior Bengali cyberbullying detection baselines by approximately 2.8 percentage points, whilst the stacking ensemble attains the highest discriminative power at 0.990 AUC, establishing a reproducible benchmark for fine-grained Bengali abuse detection.

References
  • T. T. Khan, A. Hassan, M. F. Ahamed, and S. Islam, “Multilabel bengali abusive comments classification using problem transformation method,” in 2023 20th International Conference on Electrical Engineering, Computing Science and Automatic Control, CCE 2023, Institute of Electrical and Electronics Engineers Inc., 2023.
  • A. Goni, M. U. F. Jahangir, R. R. Chowdhury, F. Akter, and K. Hussain, “Detection of suicidal ideation on social media using machine learning approaches,” International Journal of Computer Applications, vol. 186, pp. 1–8, November 2024.
  • M. T. Ahmed, M. Rahman, S. Nur, A. Z. M. T. Islam, and D. Das, “Natural language processing and machine learning based cyberbullying detection for bangla and romanized bangla texts,” TELKOMNIKA Telecommunication Computing Electronics and Control, vol. 20, no. 1, pp. 89–97, 2022.
  • G. I. Hamza, T. Muthaki, S. I. Masuk, M. M. Rahman, S. K. S. Joy, and F. M. Shah, “Bengali cyberbullying text detection: A comprehensive study,” in 2023 26th International Conference on Computer and Information Technology, ICCIT 2023, Institute of Electrical and Electronics Engineers Inc., 2023.
  • M. Sarker, M. F. Hossain, F. R. Liza, S. N. Sakib, and A. A. Farooq, “A machine learning approach to classify anti-social bengali comments on social media,” in 2022 International Conference on Advancement in Electrical and Electronic Engineering, ICAEEE 2022, Institute of Electrical and Electronics Engineers Inc., 2022.
  • M. E. Islam, S. Hossain, L. Chowdhury, M. S. Hossain, and N. Mohammed, “SentiGOLD: A large bangla gold standard multi-domain sentiment analysis dataset and its evaluation,” in Proceedings of the ACM International Conference on Proceeding Series, ACM, 2023. Funded by Bangladesh Computer Council (BCC) under EBLICT program.
  • R. Haque, N. Islam, M. Tasneem, and A. K. Das, “Multi-class sentiment classification on bengali social media comments using machine learning,” International Journal of Cognitive Computing in Engineering, vol. 4, pp. 21–35, 6 2023.
  • N. R. Bhowmik, M. Arifuzzaman, and M. R. H. Mondal, “Sentiment analysis on bangla text using extended lexicon dictionary and deep learning algorithms,” Array, vol. 13, p. 100123, 2022.
  • A. Barua, S. J. Ahmed, and M. A. A. Ansary, “A hybrid banglabert-ensemble framework for cyberbullying detection in bengali and chattogram dialects,” in 2026 IEEE 2nd International Conference on Quantum Photonics, Artificial Intelligence, and Networking (QPAIN), IEEE, 2026.
  • A. A. Miazee and A. Goni, “Agrichatnet: A dual-branch cnn with bengali chatbot integration for tomato leaf disease diagnosis,” in 2026 5th International Conference on Electrical, Computer & Telecommunication Engineering (ICECTE), 2026.
  • S. S. Nath, R. Karim, and M. H. Miraz, “Deep learning based cyberbullying detection in bangla language,” Annals of Emerging Technologies in Computing (AETIC), vol. 8, no. 1, pp. 1–15, 2024.
  • S. Sifath, T. Islam, M. Erfan, S. K. Dey, M. M. U. Islam, M. Samsuddoha, and T. Rahman, “Recurrent neural network based multiclass cyber bullying classification,” Natural Language Processing Journal, vol. 9, p. 100111, 12 2024.
  • H. Mahmud, H. Mahmud, and M. R. A. Rashid, “Enhancing sentiment analysis in bengali texts: A hybrid approach using lexicon-based algorithm and pretrained language model Bangla-BERT,” arXiv preprint arXiv:2411.19584v2, 2025.
  • K. M. M. Uddin, H. Hamim, M. N. T. Mim, A. Akhter, and M. A. Uddin, “Machine learning and deep learning-based approach to categorize bengali comments on social networks using fused dataset,” PLoS ONE, vol. 19, no. 10, p. e0308862, 2024.
  • T. Mahmud and A. C. Roy, “A multimodal multiclass cyberbullying classification in romanized bangla social media content,” in Proceedings of the 2025 28th International Conference on Computer and Information Technology (ICCIT), pp. 1–6, IEEE, 2025.
  • S. Saha, M. S. Islam, M. M. Alam, M. M. Rahman, M. Z. H. Majumder, M. S. Alam, and M. K. Hossain, “Bengali cyberbullying detection in social media using machine learning algorithms,” in 2023 5th International Conference on Sustainable Technologies for Industry 5.0, STI 2023, Institute of Electrical and Electronics Engineers Inc., 2023.
  • S. Sihab-Us-Sakib, M. R. Rahman, M. S. A. Forhad, and M. A. Aziz, “Cyberbullying detection of resource constrained language from social media using transformer-based approach,” Natural Language Processing Journal, vol. 9, p. 100104, 12 2024.
  • M. A. Rahman, M. Begum, T. Mahmud, M. S. Hossain, and K. Andersson, “Analyzing sentiments in elearning: A comparative study of bangla and romanized bangla text using transformers,” IEEE Access, vol. 12, pp. 89150–89162, 2024.
  • A. A. Miazee, “Bridging local and contextual features with ichoa-cnn-lstm for robust textual emotion recognition,” in 2025 28th International Conference on Computer and Information Technology (ICCIT), IEEE, 2025.
  • D. Kumar, “Cyberbullying detection in hinglish text using muril and explainable ai,” arXiv preprint arXiv:2506.16066v1, 2025.
  • M. Zain, N. Hussain, A. Qasim, G. Mehak, F. Ahmad, G. Sidorov, and A. Gelbukh, “Ru-old: A comprehensive analysis of offensive language detection in roman urdu using hybrid machine learning, deep learning, and transformer models,” Algorithms, vol. 18, no. 7, p. 396, 2025.
  • A. A. Miazee and A. Goni, “Bridging local and contextual features with IChOA-CNN-LSTM for robust textual emotion recognition,” in 2025 28th International Conference on Computer and Information Technology (ICCIT), pp. 4651–4656, 2025.
  • A. Goni, “A unified neural framework for fact verification using llm-gnn-lstm hybrid model,” in 2025 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE), IEEE, 2025.
  • M. R. Faisal, “Bengali cyber bullying dataset,” 2022.
  • A. Goni, M. U. F. Jahangir, and R. R. Chowdhury, “Optimized intrusion detection in iot networks using machine learning for enhanced security,” in 2024 27th International Conference on Computer and Information Technology (ICCIT), 2024.
  • A. Goni, R. J. Momtahina, and R. R. Chowdhury, “A hybrid voting-based feature selection approach for identifying intrusions in iot networks,” in 2025 IEEE International Women in Engineering (WIE) Conference on Electrical and Computer Engineering (WIECON-ECE), 2025.
  • N. Absar and M. M. Islam, “Cyberbullying detection from bangla text using cascaded deep hybrid network,” in 2024 International Conference on Innovations in Science, Engineering and Technology: Innovative Technologies for Global Solutions, ICISET 2024, Institute of Electrical and Electronics Engineers Inc., 2024.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Bengali Cyberbullying Detection Hybrid Ensemble Learning Convolutional Neural Network Multi-Class Text Classification Natural Language Processing Low-Resource Language Processing

Powered by PhDFocusTM