TY - EJOU AU - Yao, Zheng AU - Liu, Jie AU - Deng, Changjun AU - Lin, Wang TI - Cooperative Task Offloading in Mobile Edge Computing via an Improved MASAC Framework T2 - Computers, Materials \& Continua PY - VL - IS - SN - 1546-2226 AB - Mobile edge computing (MEC) is an effective paradigm for supporting latency-sensitive and computation-intensive intelligent applications. However, in dynamic mobile-edge network scenarios, mobile terminals experience time-varying wireless links due to mobility. Tasks may also arrive unpredictably, while multiple terminals compete for limited edge resources. As a result, MEC systems may suffer from service congestion and unbalanced resource utilization, which increases end-to-end latency and energy consumption. This paper investigates cooperative task offloading in dynamic MEC networks. The considered system comprises one macro base station and multiple small base stations equipped with edge-computing resources. In each time slot, each mobile terminal selects a service option, determines the task offloading ratio, and chooses its transmit power for task uploading. This sequential decision process is formulated as a multi-agent problem with continuous action spaces. Under the centralized training and decentralized execution (CTDE) framework, the problem is further modeled as a decentralized partially observable Markov decision process (Dec-POMDP). Standard multi-agent soft actor-critic (MASAC) is not fully suitable for this problem. Its original action model does not handle bounded continuous actions well. Its exploration strength may also be unsuitable at different training stages. Frequent policy updates can further make training unstable when critic estimates are inaccurate. To address these issues, this paper develops an adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC. This method uses a Beta policy to model bounded actions. It adjusts the entropy coefficient during training and delays policy updates to reduce training oscillations. Experimental results show that, under a unified training budget and a consistent evaluation protocol, the proposed method achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings in terms of overall reward, average end-to-end latency, and average energy consumption. In the large-scale scenario, compared with MASAC, it improves the overall reward by 17.8%, reduces the average end-to-end latency by 18.0%, and lowers the average energy consumption by 11.4%. KW - Mobile edge computing; multi-agent reinforcement learning; MASAC; cooperative task offloading DO - 10.32604/cmc.2026.084892