Open Access
ARTICLE
From Binary to Multi-Class: LLM-Judged Synthetic Annotation Applied to Hate Speech Detection
Departamento de Ciencias de la Computación, Universidad de Alcalá, Alcalá de Henares, Madrid, Spain
* Corresponding Author: Eva Garcia-Lopez. Email:
Computers, Materials & Continua 2026, 89(1), 51 https://doi.org/10.32604/cmc.2026.083252
Received 31 March 2026; Accepted 11 June 2026; Issue published 13 August 2026
Abstract
The increasing prevalence of hate speech on social media platforms has spurred research aimed at mitigating this societal harm. However, the development of effective machine learning solutions is hindered by a lack of labelled hate speech data in languages beyond English, particularly when attempting granular, multi-class classification. This research aims to address this data scarcity by introducing a novel methodology leveraging the ‘Large Language Model as a judge’ paradigm to transform existing binary-labelled hate speech data into multi-class datasets. Our approach aims to generate balanced datasets and enables classification across seven identity groups: race, religion, origin, gender, sexuality, age, and disability. The methodology has been applied to the Spanish Hate Speech Superset, and it has been validated using the Measuring Hate Speech dataset, demonstrating significant efficacy and broad applicability. Specifically, our approach obtains a higher match rate with human labels and a lower number of mismatches when compared with prompt-only strategies. In addition, Cohen’s Kappa scores demonstrate that our approach exhibits a moderate strength of agreement with human annotators, outperforming prompt-only strategies’ scores by 5%. As a result of the application of the proposed strategy to the Spanish Hate Speech Superset dataset, a multi-class version is obtained, comprising 6325 hateful samples classified among the seven identity groups. The proposed strategy offers a versatile solution for nuanced classification tasks beyond hate speech, providing a valuable technique for detailed categorisation in various domains.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools