Toward Secure and Adaptive Medical Digital Twins: A Privacy-Preserving Federated Multi-Agent Reinforcement Learning Framework
Tallha Akram1,*, Sadiq Ahmad2,*, Meshal Alharbi3
1 Department of Information Systems, College of Computer Engineering and Sciences, Prince Sattam bin Abdulaziz University, Al-Kharj, Saudi Arabia
2 COMSATS University Islamabad, Wah Campus, Electrical Engineering Department, Wah Cantt, Pakistan
3 Department of Computer Science, College of Computer Engineering and Sciences, Prince Sattam bin Abdulaziz University, Al-Kharj, Saudi Arabia
* Corresponding Author: Tallha Akram. Email:
; Sadiq Ahmad. Email:
Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.081458
Received 02 March 2026; Accepted 12 June 2026; Published online 15 July 2026
Abstract
Scalability limitations, privacy risks, and lack of adaptability remain key challenges in centralized medical digital win (MDT) architectures. While federated learning (FL) mitigates the need to share raw data, it often lacks adaptability to dynamic clinical environments and does not fully integrate formal privacy guarantees into the learning process. To address these challenges, this paper proposes a decentralized, federated, multi-agent reinforcement learning (F-MARL) framework to coordinate MDTs in the presence of partial observability. The framework is formulated as a multi-agent partially observable Markov decision process (MA-POMDP), enabling distributed policy optimization in heterogeneous and uncertain clinical settings. We introduce a novel algorithm, privacy-aware advantage actor–Critic with personalization and privacy protection (PA3C-PP), which integrates (i) differentially private gradient perturbation, (ii) weighted federated aggregation, and (iii) adaptive global–local policy fusion for personalization. Unlike many existing healthcare-oriented federated learning approaches that treat privacy mainly as an external or post-optimization mechanism, the proposed method incorporates differential privacy within the reinforcement learning update process. Experimental evaluation in a smart-ICU simulation demonstrates improved learning stability, enhanced resilience to agent failures, and stronger privacy protection, while maintaining latency and operational costs comparable to non-private federated baselines. These findings indicate that privacy-aware federated reinforcement learning provides a promising direction for scalable, adaptive, and regulation-compliant decentralized healthcare intelligence.
Keywords
Medical digital twin; federated multi-agent reinforcement learning; PA3C-PP; MA-POMDP