Abstract
Uncrewed Aerial Vehicle (UAV)-assisted wireless networks are emerging as a key enabler for next-generation connectivity. Under stringent Energy Efficiency (EE) requirements and finite-blocklength reliability constraints, achieving timely information delivery in UAV-assisted systems becomes particularly challenging. In this paper, we investigate a Short-Packet Communication (SPC)-enabled distributed multi-UAV system and aim to minimize the long-term average Age of Information (AoI) under finite-blocklength reliability constraints and EE requirements. To address the coordination challenge among distributed UAVs, we first develop a Graph Neural Network (GNN)-enhanced Multi-Agent Deep Reinforcement Learning (MADRL) architecture that explicitly captures inter-UAV interactions and enables scalable cooperative decision-making. Building upon this, we further propose a hierarchical reinforcement learning framework with two distinct time scales, where an upper-layer Multi-Agent Soft Actor-Critic (MASAC) optimizes continuous UAV 3D trajectories, while a lower-layer Multi-Agent Proximal Policy Optimization (MAPPO) coordinates discrete communication decisions, including transmit power control and adaptive cluster cooperation (cluster zooming). Simulation results show that the proposed H-GMSP scheme reduces the average AoI by up to 40.5% over the non-hierarchical MAPPO baseline and by up to 39.1% over the hierarchical MADQN scheme, while the joint optimization of trajectory, transmit power, and cluster scaling further achieves 26%–34% AoI reduction compared with ablated variants across diverse network scenarios.