arrow
Return

Event-Aware Instructed Assistant for Referring Video Segmentation

delete2026-06-02
delete0
PRE
AI
J
Jinyu Liu
H
Henghui Ding
S
Shuting He
Y
Yu–Gang Jiang
DOI:10.1109/TIP.2026.3697653delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact that a video typically contains multiple distinct events. Under such a mechanism, the model needs to directly understand all the complex content in the video and text, which can easily lead to confusion and hallucinations. To address this issue, we propose to decompose a video to a set of simple events by learnable Event Query, and understand complex video content in an event-by-event, easy-to-understand manner. This is based on the observation that natural language expressions often divide a video into distinct, text-related segments, each representing a separate event within a compound event. We introduce EVIS, an Event-Aware Video Instructed Segmentation Assistant, which utilizes text-guided Event Queries to partition a video into simple events, extracting event-aware visual-text features to achieve a hierarchical understanding of the video. Additionally, we propose Object-Pixel-Hybrid Learning, which enables the MLLMs to track targets in long-term videos by integrating fine-grained pixel features with prior object queries. Extensive experimental results on 5 public benchmarks demonstrate EVIS’s strong performance in addressing the referring video segmentation task. Code and trained models will be publicly released.
Keywords:
Referring video object segmentation
event-aware
multi-modal learning

Journal

IEEE Transactions on Image Processing cover
IEEE Transactions on Image Processing
IF:
13.7
Papers:
1.0W
Citations:
8.4W

Organization

S
shanghai university of finance and economics
Scholars:
235
Papers: 180
Citations: 4
F
fudan university
Scholars:
11.6W
Papers: 7.7W
Citations: 121