Abstract
•Comprehensive survey of synergistic LiDAR-Camera fusion for 3D detection.•Taxonomy of distillation as a mechanism for latent multi-modal feature fusion.•Critical analysis of cross-modal fusion via heterogeneous knowledge transfer.•Roadmap for efficient perception through multi-to-single modal feature fusion.
Recent advances in autonomous driving have greatly promoted the development of 3D object detection. Nevertheless, state-of-the-art 3D detectors usually require large-scale networks and heavy computational costs, which severely restrict their deployment on resource-limited onboard platforms. Knowledge distillation (KD), which transfers knowledge from a high-capacity teacher model to a compact student model, has become an effective solution for improving the efficiency of 3D detection while maintaining high accuracy. This paper presents a comprehensive review of knowledge distillation methods for 3D object detection in autonomous driving. We first summarize the mainstream paradigms of 3D object detection, including LiDAR-based, camera-based, and multi-modal fusion frameworks, together with the fundamental categories of knowledge distillation. Then, existing KD-based 3D detection methods are systematically reviewed from three representative perspectives: LiDAR-to-LiDAR, LiDAR-to-camera, and multi-modal-to-single-modal distillation. The characteristics, advantages, and limitations of different distillation strategies are analyzed and compared in detail. Finally, we discuss the major challenges and open issues in current KD-based 3D object detection, and outline several promising directions for future research. This survey aims to provide a concise yet comprehensive reference for the design and deployment of efficient 3D perception systems in autonomous driving. We also provide a continuously updated project homepage to support ongoing research in this area here.