arrow
Return

SGNet: Sequence Grouping Network via Vision-Language Model for Text-Guided Video Summarization

delete2025-10-27
delete0
PRE
AI
J
Jiacheng Yao
张静 cover
张静 (Jing Zhang)
卓力 cover
卓力 (Zhuo Li)
DOI:10.1109/JSTSP.2025.3626259delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video summarization can condense the key information in a video into a concise format to help viewers quickly grasp the core content. Existing approaches typically use a single model to generate uniform summaries, ignoring diverse subjective human cognition and further widening the semantic gap. Recently, multimodal large language models (MLLMs) have shown great promise in bridging this gap by leveraging the advanced semantic analysis capabilities of large language models (LLMs) and by generating high-level concepts that are closely aligned with human cognition. As an important subset of MLLMs, vision-language models (VLMs) focus on alignment and semantics extraction of visual and textual modalities for video summarization. Inspired by human cognition, we propose a sequence grouping network (SGNet) via VLM pipeline to generate text-guided video summaries aligned with high-level concepts. Firstly, a backbone with low-rank and sparsity constraints is utilized to extract frame-level features with high value from the spatial dimension. Then, sequences of frame-level features are grouped along the temporal dimension and processed using a sequence Transformer-in-Transformer (S-TNT) to model inter-frame correlations and generate compact spatio-temporal representations. Finally, the text features are obtained from video text description in the VLM pipeline to guide the cross-modal attention mechanism to embed textual information into the S-TNT, and then the vision-language knowledge transfer (VL-KT) is used to seamlessly integrate the visual and textual information into the vision-language feature to generate the summary of the video. Empirical evaluations demonstrate that the proposed SGNet achieves state-of-the-art performance on four publicly available datasets (TVSum, SumMe, UT Ego, and VaTeX), attaining F1-Scores of 67.57%, 55.80%, 62.80%, and 52.46%, respectively, with an inference speed of 5.84 FPS.
Keywords:
Sequence grouping network
vision-language model (VLM)
text-guided
video summarization

Journal

IEEE Journal of Selected Topics in Signal Processing cover
IEEE Journal of Selected Topics in Signal Processing
IF:
13.7
Papers:
1.9K
Citations:
1.1W

Organization

B
Beijing University of Technology
Scholars:
2.8W
Papers: 2.1W
Citations: 2.7W