Abstract
Text-video retrieval (TVR), an important branch of multimodal learning, has achieved significant progress driven by methods using large-scale pre-trained models. However, the general training strategy in TVR usually treats all text-video pairs equally, disregarding the unequal contributions of different pairs to the learning process. This training strategy obscures the "focus content" and limits the overall performance. Meanwhile, hard negative sample pairs result in semantically similar pairs being treated as negative samples, leading to incorrect supervision that adversely affects the search performance. To address these issues, we propose a Re presentation Re construction with S elf- P aced Learning method (Re ² SP) for the TVR task. This method integrates a self-paced learning strategy combined with the negative sample selection scheme, simulating the human cognitive process. By dynamically prioritizing high-payoff sample pairs, Re ² SP reduced the adverse effects caused by hard negative samples. To be specific, we introduce a dynamic screening scheme to filter hard negative samples and remove their influence during training, preventing false supervision on the model. Based on the current training stage, we dynamically adjust the weight of each sample, enabling the model to progressively learn from simple to complex cases. Moreover, to bridge the inherent heterogeneity gap between video and text, we develop a feature reconstruction module based on cross-modal interaction, which constructs homogeneous representations across modalities and effectively reduces the modal discrepancy. Experimental results show that our method achieves significant improvements over the baseline and establishes state-of-the-art performance on several benchmark datasets, including MSR-VTT, DiDeMo, LSMDC, and MSVD. The source code is available at https://github.com/rzheng77/Re2SP-text-video-retrieval .