Return
Topic-aware video summarization using multimodal transformer
DOI:10.1016/j.patcog.2023.109578.png)
Abstract
En 中文
Video summarization aims to generate a short and compact summary to represent the original video. Existing methods mainly focus on how to extract a general objective synopsis that precisely summaries the video content. However, in real scenarios, a video usually contains rich content with multiple top-ics and people may cast diverse interests on the visual contents even for the same video. In this pa -per, we propose a novel topic-aware video summarization task that generates multiple video summaries with different topics. To support the study of this new task, we first build a video benchmark dataset by collecting videos from various types of movies and annotate them with topic labels and frame-level importance scores. Then we propose a multimodal Transformer model for the topic-aware video summa-rization, which simultaneously predicts topic labels and generates topic-related summaries by adaptively fusing multimodal features extracted from the video. Experimental results show the effectiveness of our method. (c) 2023 Elsevier Ltd. All rights reserved.
Keywords:
Topic-aware video summarization
Multimodal transformer
Video summarization dataset
Journal
IF:
7.6
Papers:
1.3W
Citations:
4.5W

