TY - EJOU AU - Moreno-Cediel, Antonio AU - Garcia-Cabot, Antonio AU - Garcia-Lopez, Eva TI - From Binary to Multi-Class: LLM-Judged Synthetic Annotation Applied to Hate Speech Detection T2 - Computers, Materials \& Continua PY - VL - IS - SN - 1546-2226 AB - The increasing prevalence of hate speech on social media platforms has spurred research aimed at mitigating this societal harm. However, the development of effective machine learning solutions is hindered by a lack of labelled hate speech data in languages beyond English, particularly when attempting granular, multi-class classification. This research aims to address this data scarcity by introducing a novel methodology leveraging the ‘Large Language Model as a judge’ paradigm to transform existing binary-labelled hate speech data into multi-class datasets. Our approach aims to generate balanced datasets and enables classification across seven identity groups: race, religion, origin, gender, sexuality, age, and disability. The methodology has been applied to the Spanish Hate Speech Superset, and it has been validated using the Measuring Hate Speech dataset, demonstrating significant efficacy and broad applicability. Specifically, our approach obtains a higher match rate with human labels and a lower number of mismatches when compared with prompt-only strategies. In addition, Cohen’s Kappa scores demonstrate that our approach exhibits a moderate strength of agreement with human annotators, outperforming prompt-only strategies’ scores by 5%. As a result of the application of the proposed strategy to the Spanish Hate Speech Superset dataset, a multi-class version is obtained, comprising 6325 hateful samples classified among the seven identity groups. The proposed strategy offers a versatile solution for nuanced classification tasks beyond hate speech, providing a valuable technique for detailed categorisation in various domains. KW - Hate speech; synthetic data annotation; dataset; LLM-as-a-judge; multi-class classification DO - 10.32604/cmc.2026.083252