返回
P2M: Progressive Perspective Mining for ReferringVideo Object Segmentation
DOI:10.1109/TMM.2025.3618539.png)
摘要
En 中文
Referring video object segmentation (RVOS) aims to segment the object instances referred to by linguistic expressions in video frames. The prevailing approaches mainly rely on simplistic fusion strategies, wherein textual features are directly interacted with video features without considering the impact of textual semantics at different levels. These coarse-grained fusion strategies hinder the model's ability to perceive changes in object appearance and movement, resulting in performance degradation. To mitigate this issue, we propose a Progressive Perspective Mining (P$<^>{2}$M) framework, which leverages a coarse-to-fine perspective to mine latent information from text and video, enabling precise segmentation of referred objects. P$<^>{2}$M consists of two key components: Progressive Vision-Language Interaction (PVLI) and Vision-Language Synergistic Fusion (VLSF). Specifically, PVLI leverages language features across subject, word, and sentence levels to mine textual information, enabling a progressive interaction with video features within an integrated representational space. Concurrently, VLSF focuses on generating semantically rich object queries for segmentation by employing slot attention mechanisms to mine and integrate relevant visual features with linguistic semantics. Furthermore, we introduce two query optimization losses: (1) the Matching Optimization Loss constrains the best queries between frame-level and video-level, effectively preventing the queries of the tracking target from drifting along the temporal dimension during the inference phase; (2) the Vision-Language Semantic Alignment Loss performs a word-by-word matching between object queries and expression, aligning the multi-modal joint space and enhancing the framework's understanding of the textual description. We conducted various experiments on the RVOS task, achieving new state-of-the-art results across all benchmarks, thereby demonstrating the effectiveness of P$<^>{2}$M.
Keyword:
Videos
Feature extraction
Semantics
Linguistics
Transformers
Visualization
Object segmentation
Image segmentation
Query processing
Optimization
Coarse-to-fine
progressive interaction
slot attention
video object segmentation
期刊
IF:
9.7
论文数:
4.5K
被引数:
2.4W

