Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.086535
Special Issues
Table of Content

Open Access

ARTICLE

An Improved Safe Soft Actor-Critic Path Planning Algorithm for Autonomous Vehicles Based on a Dual-Stream Q-Network and Dynamic Analytic Hierarchy Process

Shengxuan Dong, Xiongwei Li*
Shijiazhuang Campus, Army Engineering University of PLA, Shijiazhuang, China
* Corresponding Author: Xiongwei Li. Email: email

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.086535

Received 01 June 2026; Accepted 21 July 2026; Published online 12 August 2026

Abstract

To address the conflict between navigation performance and safety constraints in safe reinforcement learning, this paper proposes Dual Stream-Analytic Hierarchy Process-Safe Soft Actor (DS-AHP-SAC), a safe soft actor-critic algorithm based on a dual-stream Q-network and dynamic Analytic Hierarchy Process (AHP) stratified experience replay. The algorithm achieves a balance between reward maximization and constraint satisfaction through three synergistic designs: (1) decoupling the Q-network into independent navigation and safety value streams to eliminate gradient interference at the Critic level and mitigate gradient competition at the Actor level; (2) constructing a three-criterion dynamic sampling strategy based on AHP, incorporating safety urgency, information value, and scarcity to enable phase-adaptive experience replay; (3) designing a curriculum-scheduling scheme that linearly increases the safety constraint weight during training, preventing policy degradation caused by premature imposition of high safety penalties. Experimental results in a two-dimensional continuous navigation environment demonstrate that DS-AHP-SAC reduces the violation rate by 20.0% compared to unconstrained SAC without sacrificing navigation success rate, while avoiding the policy collapse observed in SAC-Lagrangian. Ablation studies validate the necessity of the dual-stream Q-network and the convergence acceleration effect of dynamic AHP.

Keywords

Safe reinforcement learning; dual-stream Q-network; analytic hierarchy process; experience replay; constrained Markov decision process
  • 137

    View

  • 32

    Download

  • 0

    Like

Share Link