Multi-UAV Collaborative Energy Charging for Battery-Free SWIPT-Enabled Sensor Networks Based on MADDPG
Xiangyi Le1, Deyu Lin1,2,*, Yufei Zhao2, Wang Miao3, Yong Liang Guan2
1 School of Software, Nanchang University, Nanchang, China
2 School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore, Singapore
3 Computer Science Department, University of Exeter, Exeter, UK
* Corresponding Author: Deyu Lin. Email:
(This article belongs to the Special Issue: Intelligent IoT for Smart Cities and Sustainable Energy Systems)
Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.083306
Received 01 April 2026; Accepted 10 June 2026; Published online 07 July 2026
Abstract
The emergence of Unmanned Aerial Vehicle (UAV)-enabled Wireless Energy Transfer (WET) and Simultaneous Wireless Information and Power Transfer (SWIPT) technology provide a promising solution to overcome the energy sustainability limitations of traditional harvesting-reliant sensor networks. However, in large-scale Battery-free SWIPT-enabled Sensor Networks (BSSN) characterized by sparse node distribution and heterogeneous energy consumption and harvesting rates, employing a single UAV for energy replenishment often suffers from insufficient operation continuity and low charging efficiency. To overcome these challenges, a Multi-UAV Collaborative Energy Charging for BSSN Based on Multi-Agent Deep Deterministic Policy Gradient (MCEC-MADDPG) is proposed in this paper. Specifically, we construct a collaborative one-to-one precision energy supply model where UAVs hover directly above specific nodes to achieve power transmission without complex beamforming requirements. To achieve collaborative scheduling among multiple UAVs in wide-area dynamic environments, the energy replenishment problem is first formulated as a Partially Observable Markov Decision Process (POMDP). Subsequently, the Centralized Training with Decentralized Execution (CTDE) architecture of the MADDPG algorithm is leveraged to solve this POMDP, which effectively tackles the non-stationarity challenge inherent in multi-agent environments. Simulation results demonstrate that MCEC-MADDPG exhibits superior performance in terms of convergence speed and stability. It enables the adaptive emergence of spatial-division collaborative strategies, significantly enhances the average residual energy of the network, and elevates the node survival rate to nearly
90% . Compared with Deep Deterministic Policy Gradient (DDPG), the traditional static Partition-Greedy method, the heuristic K-Means algorithm and the dynamic Two-Layer task allocation strategy, the proposed approach demonstrates substantial advantages.
Keywords
Battery-free SWIPT-enabled sensor networks; multi-agent deep deterministic policy gradient; multi-unmanned aerial vehicle; collaborative energy charging; partially observable Markov decision process; centralized training with decentralized execution