Return
Multi-Temporal Granularity Concept Induction for semantically driven video summarization
DOI:10.1016/j.eswa.2025.127128.png)
Abstract
En 中文
The existing video summarization methods mainly focus on extracting and analyzing visual features, but often overlook the higher-level semantic connections between video frames. This approach, while addressing surface- level visual elements, fails to fully understand the complex scenes, characters, and events of videos and their temporal associations, resulting in summaries that lack representativeness. To address this issue, this paper proposes a video summarization method using Multi-Temporal Granularity Concept Induction, MTGC-VS. This method first models the temporal connections and semantic information of the input video through a concept encoder at multiple granularities, identifying the most representative semantic prototypes in the video as video concepts. These concepts are then integrated to form the general semantics of the video. Based on this, a semantic enhancer strengthens the relevance of each frame to general semantics, identifying the key content that aligns best with the general meaning, enhancing the semantic consistency between the summary and the original video, and making the summary more semantically representative. The performance of the proposed method is validated through extensive experiments conducted on the TVSum and SumMe datasets and compared with that of the current state-of-the-art methods. The experimental results demonstrate the effectiveness of the MTGC-VS method.
Keywords:
Video summarization
Multi-temporal granularity
Concept induction
General semantic
Journal
IF:
7.5
Papers:
2.9W
Citations:
10.2W
Organization
No organization information available

