arrow
Return

Evaluating SAM2 for Video Semantic Segmentation

delete2026-08-24
delete0
PRE
AI
S
Syed Ariff Syed Hesham
Y
Yun Liu *
G
Guolei Sun *
J
Jing Yang
H
Henghui Ding
X
Xue Geng
蒋旭东 cover
蒋旭东 (Xudong Jiang)
DOI:10.1007/s11633-026-1638-9delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
The segment anything model 2 (SAM2) has proven to be a powerful foundation model for promptable visual object segmentation in both images and videos, capable of storing object-aware memories and transferring them temporally through memory blocks. While SAM2 excels in video object segmentation by providing dense segmentation masks based on prompts, extending it to dense video semantic segmentation (VSS) poses challenges due to the need for spatial accuracy, temporal consistency, and the ability to track multiple objects with complex boundaries and varying scales. This paper explores the extension of SAM2 for VSS, focusing on two primary approaches and highlighting firsthand observations and common challenges faced during this process. The first approach involves using SAM2 to extract unique objects as masks from a given image, with a segmentation network employed in parallel to generate and refine initial predictions. The second approach utilizes the predicted masks to extract unique feature vectors, which are then fed into a simple network for classification. The resulting classifications and masks are subsequently combined to produce the final segmentation. Our experiments suggest that leveraging SAM2 enhances overall performance in VSS, primarily due to its precise predictions of object boundaries.
Keywords:
Video semantic segmentation (VSS)
segment anything model 2 (SAM 2)
visual foundation models
boundary refinement
temporal consistency
mask-based classification

Journal

Machine Intelligence Research cover
Machine Intelligence Research
IF:
8.7
Papers:
301
Citations:
882

Organization

S
State Key Laboratory of Public Big Data
Scholars:
18
Papers: 11
Citations: 0
S
School of Electrical and Electronic Engineering
Scholars:
267
Papers: 114
Citations: 3
F
fudan vision and learning lab
Scholars:
2
Papers: 1
Citations: 0
I
Institute for Infocomm Research
Scholars:
46
Papers: 43
Citations: 662
researcher View more organizations