arrow
返回

P2M: Progressive Perspective Mining for ReferringVideo Object Segmentation

delete2025-01-01
delete0
PRE
AI
Y
Yihan Wang
B
Baoli Sun
X
Xinzhu Ma
葛
葛宏伟 (Hongwei Ge) *
F
Fan, Jiulin
H
Haojie Li
DOI:10.1109/TMM.2025.3618539delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Referring video object segmentation (RVOS) aims to segment the object instances referred to by linguistic expressions in video frames. The prevailing approaches mainly rely on simplistic fusion strategies, wherein textual features are directly interacted with video features without considering the impact of textual semantics at different levels. These coarse-grained fusion strategies hinder the model's ability to perceive changes in object appearance and movement, resulting in performance degradation. To mitigate this issue, we propose a Progressive Perspective Mining (P$<^>{2}$M) framework, which leverages a coarse-to-fine perspective to mine latent information from text and video, enabling precise segmentation of referred objects. P$<^>{2}$M consists of two key components: Progressive Vision-Language Interaction (PVLI) and Vision-Language Synergistic Fusion (VLSF). Specifically, PVLI leverages language features across subject, word, and sentence levels to mine textual information, enabling a progressive interaction with video features within an integrated representational space. Concurrently, VLSF focuses on generating semantically rich object queries for segmentation by employing slot attention mechanisms to mine and integrate relevant visual features with linguistic semantics. Furthermore, we introduce two query optimization losses: (1) the Matching Optimization Loss constrains the best queries between frame-level and video-level, effectively preventing the queries of the tracking target from drifting along the temporal dimension during the inference phase; (2) the Vision-Language Semantic Alignment Loss performs a word-by-word matching between object queries and expression, aligning the multi-modal joint space and enhancing the framework's understanding of the textual description. We conducted various experiments on the RVOS task, achieving new state-of-the-art results across all benchmarks, thereby demonstrating the effectiveness of P$<^>{2}$M.
Keyword:
Videos
Feature extraction
Semantics
Linguistics
Transformers
Visualization
Object segmentation
Image segmentation
Query processing
Optimization
Coarse-to-fine
progressive interaction
slot attention
video object segmentation

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

B
beihang university
学者数:
5.2K
论文数: 2.0K
被引数: 21
D
Dalian University of Technology
学者数:
6.0W
论文数: 4.4W
被引数: 5.5W
引用论文

引用论文

Segmentation from Natural Language Expressions
err2016-09-17
err0
errOAAI
errRonghang Hu; Marcus Rohrbach; Trevor Darrell
err分享
err收藏
Actor and Action Modular Network for Text-Based Video Segmentation
err2022-01-01
err0
errOAAI
errJianhua Yang; Yan Huang; Kai Niu; Linjiang Huang; Zhanyu Ma; Liang Wang
err分享
err收藏
学者 查看更多内容