Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.084177
Special Issues
Table of Content

Open Access

ARTICLE

Automated Hate Speech Profiling via Lexicon-Enriched Ensemble Learning and Ego-Network Analysis

Sayfudin Sayfudin1,2, Deris Stiawan3,*, Ferdiansyah Ferdiansyah4, Rahmat Budiarto5
1 Faculty of Engineering, Universitas Sriwijaya, Palembang, Indonesia
2 Bureau of Data and Information, Universitas Muhammadiyah Palembang, Palembang, Indonesia
3 Faculty of Computer Science, Universitas Sriwijaya, Palembang, Indonesia
4 Faculty of Computer Science, Universitas Indo Global Mandiri, Palembang, Indonesia
5 College of Computing and Information, Al-Baha University, Al Aqiq, Saudi Arabia
* Corresponding Author: Deris Stiawan. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.084177

Received 17 April 2026; Accepted 01 June 2026; Published online 22 July 2026

Abstract

Tightening global regulation of digital toxicity demands hate-speech detection that is accurate, explainable, traceable, and forensically usable. The challenge intensifies in multilingual and code-mixed settings such as Indonesian social media, where linguistic variation and informal expressions cause feature sparsity and reduce machine learning (ML) effectiveness. Most prior work emphasizes text classification while neglecting actor profiling and the network structures through which hate speech propagates. We propose Dynamic Lexicon-Driven Network (DyLex-Net), an integrated framework for profiling actors who disseminate hate speech, combining dataset-driven dynamic-lexicon analysis, classical ML ensemble validation, and ego-network analysis under a forensic-readiness orientation. The lexicon is built from a large multilingual corpus and serves as a transparent, auditable knowledge base for real-time inference. Logistic Regression (LR), Linear Support Vector Machine (SVM), and a voting ensemble are used for offline benchmarking. Experiments cover an integrated corpus of more than 715,000 posts plus real-time account-level inference. The framework achieves consistent F1 across models, with ensembles most stable. DyLex-Net produces explainable, traceable actor risk profiles that satisfy both analytical accuracy and forensic interpretability, bridging technical performance and legal requirements for multilingual hate-speech analysis and contributing to cyber threat intelligence and digital forensics.

Keywords

Hate speech detection; cyber threat intelligence; actor profiling; forensic readiness; dynamic lexicon; ego-network analysis; code-mixed text; ensemble learning
  • 187

    View

  • 33

    Download

  • 1

    Like

Share Link