Multi-Agent Deep Reinforcement Learning for UAV Swarms in Post-Disaster Relief Tasks with Dynamic Heterogeneous Action Spaces
Li, Wenxin ; Feng, Yongxin ; Zhou, Fan ; Zhang, Peiying ; Kostromitin, Konstantin Igorevich ; Guizani, Mohsen
Li, Wenxin
Feng, Yongxin
Zhou, Fan
Zhang, Peiying
Kostromitin, Konstantin Igorevich
Guizani, Mohsen
Supervisor
Department
Machine Learning
Embargo End Date
Type
Journal article
Date
License
Language
Collections
Research Projects
Organizational Units
Journal Issue
Abstract
Unmanned aerial vehicle (UAV) swarms exhibit distributed sensing and flexible formation capabilities, which facilitate sustained operations and rapid responses in high-risk, large-area tasks. As a result, they are widely used in disaster relief and related applications. In this paper, we focus on the autonomous decision-making of heterogeneous UAV swarms for accomplishing previously untrained tasks in complex, dynamically evolving post-disaster environments. The swarm's action space typically consists of heterogeneous sub-action spaces corresponding to distinct UAV functions and task modalities, which renders direct policy transfer unreliable or inefficient. Therefore, we propose a multi-agent deep reinforcement learning method for action spaces that are heterogeneous and temporally dynamic. First, a multi-round variational autoencoder is introduced to provide a unified representation of all sub-action spaces within heterogeneous UAV swarms, which mitigates heterogeneity at the representation level. Then, a graph attention network is employed to capture correlations between the represented sub-action spaces, enabling policy transfer to unseen action spaces. However, we find that environmental nonstationarity renders reward signals markedly sparse and delayed, which makes credit assignment difficult. Consequently, we introduce a dynamic hindsight experience replay mechanism that relabels failed trajectories, converting hard-to-use negative samples into informative pseudo-successful samples and thereby significantly improving sample efficiency and training stability. Simulation results in disaster relief scenarios demonstrate that the proposed algorithm generalizes learned policies robustly to unseen action spaces in environments with varying complexities, while exhibiting enhanced resilience under dynamically evolving interference conditions.
Citation
W. Li, Y. Feng, F. Zhou, P. Zhang, K.I. Kostromitin, M. Guizani, "Multi-Agent Deep Reinforcement Learning for UAV Swarms in Post-Disaster Relief Tasks with Dynamic Heterogeneous Action Spaces," IEEE Transactions on Vehicular Technology, pp. 1-16, 2026, https://doi.org/10.1109/tvt.2026.3701136.
Source
IEEE Transactions on Vehicular Technology
Conference
Keywords
46 Information and Computing Sciences, 4602 Artificial Intelligence, 4611 Machine Learning, 11 Sustainable Cities and Communities
Subjects
Source
Publisher
IEEE
