Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.086261
Special Issues
Table of Content

Open Access

ARTICLE

SD-KRE: A Method for Structural Decoupling and Knowledge Reuse Evolution of Reinforcement Learning Reward Functions Assisted by Large Language Models

Yuqing Cao, Xiliang Chen*, Legui Zhang*, Jun Lai, Haoyang Dong, Xuefei Sun, Xiaoyan Wang
College of Command and Control Engineering, Army Engineering University of PLA, Nanjing, China
* Corresponding Author: Xiliang Chen. Email: email; Legui Zhang. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.086261

Received 27 May 2026; Accepted 16 July 2026; Published online 05 August 2026

Abstract

The design of reward functions is crucial to the success of reinforcement learning, yet the process often relies on expert experience and is difficult to debug. Although large language models (LLMs) offer new opportunities for automated reward design, existing methods still face challenges such as poor interpretability, inability to reuse knowledge, and optimization blindness. To address these issues, this paper proposes a method for structural decoupling and knowledge reuse evolution, referred to as SD-KRE. Its core lies in treating the reward function as a composition of multiple structured units with clear semantics and functionally decoupled components, and in measuring the importance of each unit’s utility through a multi-dimensional quantitative evaluation method; The quantified structural units are stored in a structured reward prior knowledge base and serve as prior knowledge to guide the LLM in targeted reuse and evolutionary optimization during subsequent iterations. Experiments on six continuous control and dexterous manipulation benchmarks show SD-KRE significantly outperforms Eureka and Text2Reward in final performance, convergence speed, and stability, while matching or surpassing manually designed rewards on most tasks. Its stronger alignment with human prior knowledge further validates the semantic validity of the approach. By enabling structured reuse and directed evolution, SD-KRE overcomes the blindness and forgetfulness of conventional LLM-based reward design. It delivers expert-level rewards with higher efficiency, providing a reliable plug-and-play solution for complex reinforcement learning tasks.

Keywords

Reinforcement learning; large language models; reward design; structural decoupling; knowledge reuse; evolutionary optimization
  • 86

    View

  • 13

    Download

  • 0

    Like

Share Link