Abstract
<p>Coordinating multiple heterogeneous xApps in Open Radio Access Networks (O-RAN) is challenging because real-time inter-xApp synchronization introduces communication overhead, latency, and scalability bottlenecks under live deployments. This paper proposes an attention-based Heterogeneous-Agent Proximal Policy Optimization (HAPPO) framework that coordinates four heterogeneous xApps-power control, resource-block allocation, access control, and beam selection-through a Digital Twin (DT)-enabled training pipeline. The framework combines 1) hybrid actor heads supporting mixed continuous/discrete actions; 2) task-aware self-attention for implicit coordination without predefined communication graphs; and 3) DT-based asynchronous experience collection with delay modelling, online correction, and safe rollback. We further validate that the learned attention aligns with the physical coupling between xApps (Pearson rho = 0.78 +/- 0.04 ), establishing an interpretable link between the policy and underlying network dependencies. On a VIAVI O-RAN RIC emulator across 10 random seeds, the proposed scheme attains a 91.2% aggregate QoS score and 1.05 ms inference latency, improves per-slice SLA satisfaction by up to 6.4% over strong MARL baselines (FACMAC, MAPPO, CommNet), and reaches the 80%-of-final performance level 1.49 & times; faster than the attention-free counterpart. Improvements are statistically significant ( p < 0.05 , paired two-sided Welch t-tests with Holm-Bonferroni correction across baselines; under fluctuating UE demands and diverse service-level requirements).</p>