Abstract
In multi-modal vision studies, depth channels have complemented the representation capacity of RGB signals with extended spatial awareness, which has led to dedicated research in the RGB-D tracking field. In general, to reflect the information conveyed by the depth modality, the most essential topic for RGB-D tracking is how to perform multi-modal fusion. Current solutions employ elaborate fusion modules to model the interactions between RGB and depth channels in a fixed manner. These approaches tend to default to involving the depth modality in tracking every single frame. However, the utility and complementarity of the depth channel with respect to RGB vary across scenarios. To address this, we propose a modality informativeness controlled RGB-D tracking (MICTracker) framework that enables adaptive fusion via performing depth data quality assessment. Specifically, we design a tracking gain prediction framework to measure the depth channel quality given an RGB-D image pair. To enable the training of our framework, we measure the difference between the losses obtained by forward-passing RGB-only and RGB-D inputs through our RGB-D tracker. After quantifying such difference in terms of tracking gain, we can train our predictor, which then controls the fusion of the depth channel with RGB signals in the test phase. Explicitly equipped with the tracking gain prediction framework, an evidential RGB-D tracker is constructed with improved interpretability when performing RGB-D fusion. Extensive experimental results on standard benchmarking datasets, which outperform the state-of-the-art RGB-D trackers, demonstrate the merits of our MICTracker.