TY - EJOU
AU - Ma, Rong
AU - Ren, Jin
AU - Shen, Shaobing
AU - Li, Yunhe
AU - Hu, Man
TI - Toward Trustworthy Chinese Large Language Models: A Multi-Dimensional Evaluation of Toxicity, Bias, and Robustness
T2 - Computers, Materials \& Continua
PY -
VL -
IS -
SN - 1546-2226
AB - Large language models (LLMs) have emerged as a transformative foundation across natural language processing and intelligent systems, yet their security, robustness, and responsible deployment remain critical open challenges. In particular, the multi-dimensional evaluation of toxicity and bias in Chinese LLMs remains limited, posing significant risks for real-world applications that demand trustworthy AI. In this paper, we propose TrustEval, a dataset- and model-agnostic evaluation framework that provides a systematic assessment of Chinese LLMs from the perspectives of toxicity, bias, and robustness. Unlike existing benchmarks that focus primarily on capability, TrustEval explicitly targets model security and reliability by probing three dimensions: (1) Toxicity Measurement, which quantifies how toxic or non-toxic prompts trigger harmful model outputs; (2) Bias Measurement, which evaluates whether LLMs exhibit discriminatory behavior across sensitive attributes such as gender, race, and region; and (3) Avoidance Rate Measurement, which assesses the model’s ability to recognize and refuse toxic inputs. Experimental results on nine Chinese LLMs and two general-purpose comparison models across three datasets show that the evaluated models can generate toxic content under the tested prompting conditions and exhibit non-trivial attribute-level toxicity disparities. Furthermore, the built-in avoidance mechanisms of several models remain insufficient for robust safety enforcement. These findings reveal notable safety weaknesses in the evaluated Chinese LLMs and underscore the need for security-aware training and responsible AI design for large language models. All code, translated datasets, prompt templates, and evaluation scripts are publicly available at https://github.com/fenffef/TrustEval.
KW - Large language models; model security and robustness; toxicity evaluation; bias assessment; responsible AI; trustworthy AI
DO - 10.32604/cmc.2026.086288