Abstract
Vehicular millimetre-wave (mmWave) networks suffer from beam misalignment and throughput degradation due to high mobility, blockages, and interference-challenges that static beam strategies cannot address. This paper addresses adaptive beam configuration in dynamic vehicular scenarios through reinforcement learning (RL), i.e. Deep Q-Network (DQN) and Q-learning, to jointly optimise base station placement, beam direction, and beamwidth for maximum system sum-rate. Recognising that the lack of interpretability in learned policies remains a critical barrier to operator trust and practical deployment, we integrate a novel multi-agent large language model (LLM) explainability layer that processes user queries in parallel through specialised agents, hence combine numerical logs, scenario rules, and visual outputs into concise, human-understandable explanations. This unified framework demonstrates how high-performance RL optimisation and transparent explainability can be combined to build trust and facilitate practical deployment in future 5G/6G vehicular networks and Open RAN environments.