Abstract
Self-Organising Network (SON) control in disaggregated Open Radio Access Networks (O-RAN) involves hard optimisation problems, where functional split selection governs the placement of baseband processing across distributed units (O-DUs) and central units (O-CUs) under conflicting objectives and operator-defined bottom-line constraints. While reinforcement learning (RL) is widely used, existing approaches are typically monolithic, coupling multiple objectives in a single reward and making policies difficult to constrain and adapt. We propose CICA-SON, a collective-intelligence framework that decouples objective learning from policy enforcement via a pool of independently trained experts and a lightweight Deep QNetwork (DQN) orchestrator. Instantiated for O-RAN functional split selection with two experts, CICA-SON is evaluated on a realistic urban topology with time-varying traffic. Simulation results show that the experts converge to specialised behaviours, where the utilisation-focused expert achieves 55% average ODU utilisation compared to 2 5% for the computation-cost expert, while incurring higher normalised computation cost ( 0.70 vs. 0.35). The orchestrator converges to stable regimes that satisfy utilisation and computation-cost bottom-lines, reduce reconfiguration overhead, and adapt without retraining.