Return
Probabilistic Embeddings With Evidence Learning and Refinement for Text–Video Retrieval
D
Z
X
X
J
DOI:10.1109/tip.2026.3715851.png)
Abstract
En 中文
This paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the information granularity and abstraction levels between the two modalities hinder a reliable sample-level alignment. Moreover, redundant visual content, sparse textual descriptions, and temporal variability in videos introduce additional uncertainty, resulting in ambiguous matching and suboptimal performance. In this paper, we propose a novel method named Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR), which models video-text pairs as probability distributions and captures uncertainty through the evidence theory. Specifically, we perform distribution-level representation learning to resolve the semantic ambiguity of video-text pairs. To improve the alignment further, we introduce a distribution-based embedding refinement module to ameliorate the semantic consistency across modalities. The proposed PE2LR is able to pull positive sample pairs closer in the embedding space, while pushing the negative pairs apart. Comprehensive experiments on several benchmark datasets (including MSRVTT, DiDeMo, and ActivityNet Captions) demonstrate that our PE2LR achieves state-of-the-art search performance.
Keywords:
Text-video retrieval
probabilistic embedding
evidence learning
cross-modal learning
Journal
IF:
13.7
Papers:
1.0W
Citations:
8.4W
