Open Access
ARTICLE
Toward Trustworthy Chinese Large Language Models: A Multi-Dimensional Evaluation of Toxicity, Bias, and Robustness
1 School of Electronics and Electrical Engineering, Zhaoqing University, Zhaoqing, China
2 School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
* Corresponding Author: Yunhe Li. Email:
(This article belongs to the Special Issue: Large Language Models: Foundations, Advances, and Emerging Applications)
Computers, Materials & Continua 2026, 89(2), 32 https://doi.org/10.32604/cmc.2026.086288
Received 27 May 2026; Accepted 13 July 2026; Issue published 15 September 2026
Abstract
Large language models (LLMs) have emerged as a transformative foundation across natural language processing and intelligent systems, yet their security, robustness, and responsible deployment remain critical open challenges. In particular, the multi-dimensional evaluation of toxicity and bias in Chinese LLMs remains limited, posing significant risks for real-world applications that demand trustworthy AI. In this paper, we propose TrustEval, a dataset- and model-agnostic evaluation framework that provides a systematic assessment of Chinese LLMs from the perspectives of toxicity, bias, and robustness. Unlike existing benchmarks that focus primarily on capability, TrustEval explicitly targets model security and reliability by probing three dimensions: (1) Toxicity Measurement, which quantifies how toxic or non-toxic prompts trigger harmful model outputs; (2) Bias Measurement, which evaluates whether LLMs exhibit discriminatory behavior across sensitive attributes such as gender, race, and region; and (3) Avoidance Rate Measurement, which assesses the model’s ability to recognize and refuse toxic inputs. Experimental results on nine Chinese LLMs and two general-purpose comparison models across three datasets show that the evaluated models can generate toxic content under the tested prompting conditions and exhibit non-trivial attribute-level toxicity disparities. Furthermore, the built-in avoidance mechanisms of several models remain insufficient for robust safety enforcement. These findings reveal notable safety weaknesses in the evaluated Chinese LLMs and underscore the need for security-aware training and responsible AI design for large language models. All code, translated datasets, prompt templates, and evaluation scripts are publicly available at https://github.com/fenffef/TrustEval.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools