arrow
Return

Video Grounded Conversation Generation for Reference Surgical Instrument Segmentation

delete2026-01-01
delete0
PRE
AI
Y
Yihan Wang
Q
Qiao Yan
L
Lihao Liu
Y
Yuchen Yuan
X
Xiaowei Hu
L
Li, Jinpeng *
P
Pheng-Ann Heng
DOI:10.1007/978-3-032-09784-2_13delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Surgical instrument segmentation in videos is essential for computer-assisted interventions, enabling accurate tool identification during surgeries. However, current Grounded Conversation Generation (GCG) methods struggle with specifying the instrument of interest and capturing complex interactions in dynamic environments due to limitations in understanding intra-frame and inter-frame information. Here, we formulate a novel Video-GCG framework for improved reference surgical instrument segmentation, which combines visual data with context-aware textual descriptions. First, we develop a Temporal Dynamic Sampling (TDDS) strategy to enhance temporal-spatial feature extraction, solving the intra-frame problem. Then, we present a mask decoding strategy to refine segmentation outputs and reduce the impact of blurred or ambiguous visual information from the surrounding environment, tackling the inter-frame problem. Experimental results show that our method outperforms the state-of-the-art VIS-Net by 18.1% and 7.5% in mAP on the EndoVis-RS17 &18 datasets, showcasing superior performance and efficiency with fewer computational resources. Codes will be released.
Keywords:
Reference Surgical Instrument Segmentation
Video Grounded Conversation Generation
Multimodal Large Language Models

Journal

C
COLLABORATIVE INTELLIGENCE AND AUTONOMY IN IMAGE-GUIDED SURGERY, COLAS 2025
IF:
0
Papers:
16
Citations:
0

Organization

C
chinese university of hong kong
Scholars:
2.4K
Papers: 1.2K
Citations: 0
A
amazon.com
Scholars:
698
Papers: 505
Citations: 8
S
south china university of technology
Scholars:
6.7W
Papers: 5.1W
Citations: 85
researcher View more organizations