Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.086288
Special Issues
Table of Content

Open Access

ARTICLE

Toward Trustworthy Chinese Large Language Models: A Multi-Dimensional Evaluation of Toxicity, Bias, and Robustness

Rong Ma1, Jin Ren1, Shaobing Shen1, Yunhe Li1,*, Man Hu2
1 School of Electronics and Electrical Engineering, Zhaoqing University, Zhaoqing, China
2 School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
* Corresponding Author: Yunhe Li. Email: email
(This article belongs to the Special Issue: Large Language Models: Foundations, Advances, and Emerging Applications)

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.086288

Received 27 May 2026; Accepted 13 July 2026; Published online 12 August 2026

Abstract

Large language models (LLMs) have emerged as a transformative foundation across natural language processing and intelligent systems, yet their security, robustness, and responsible deployment remain critical open challenges. In particular, the multi-dimensional evaluation of toxicity and bias in Chinese LLMs remains limited, posing significant risks for real-world applications that demand trustworthy AI. In this paper, we propose TrustEval, a dataset- and model-agnostic evaluation framework that provides a systematic assessment of Chinese LLMs from the perspectives of toxicity, bias, and robustness. Unlike existing benchmarks that focus primarily on capability, TrustEval explicitly targets model security and reliability by probing three dimensions: (1) Toxicity Measurement, which quantifies how toxic or non-toxic prompts trigger harmful model outputs; (2) Bias Measurement, which evaluates whether LLMs exhibit discriminatory behavior across sensitive attributes such as gender, race, and region; and (3) Avoidance Rate Measurement, which assesses the model’s ability to recognize and refuse toxic inputs. Experimental results on nine Chinese LLMs and two general-purpose comparison models across three datasets show that the evaluated models can generate toxic content under the tested prompting conditions and exhibit non-trivial attribute-level toxicity disparities. Furthermore, the built-in avoidance mechanisms of several models remain insufficient for robust safety enforcement. These findings reveal notable safety weaknesses in the evaluated Chinese LLMs and underscore the need for security-aware training and responsible AI design for large language models. All code, translated datasets, prompt templates, and evaluation scripts are publicly available at https://github.com/fenffef/TrustEval.

Keywords

Large language models; model security and robustness; toxicity evaluation; bias assessment; responsible AI; trustworthy AI
  • 173

    View

  • 33

    Download

  • 0

    Like

Share Link