arrow
Return

Topic-aware video summarization using multimodal transformer

delete2023-08-01
delete7
PRE
AI
W
Wentian Zhao
R
Rui Hua
吴心筱 cover
吴心筱 (Xinxiao Wu) *
DOI:10.1016/j.patcog.2023.109578delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Video summarization aims to generate a short and compact summary to represent the original video. Existing methods mainly focus on how to extract a general objective synopsis that precisely summaries the video content. However, in real scenarios, a video usually contains rich content with multiple top-ics and people may cast diverse interests on the visual contents even for the same video. In this pa -per, we propose a novel topic-aware video summarization task that generates multiple video summaries with different topics. To support the study of this new task, we first build a video benchmark dataset by collecting videos from various types of movies and annotate them with topic labels and frame-level importance scores. Then we propose a multimodal Transformer model for the topic-aware video summa-rization, which simultaneously predicts topic labels and generates topic-related summaries by adaptively fusing multimodal features extracted from the video. Experimental results show the effectiveness of our method. (c) 2023 Elsevier Ltd. All rights reserved.
Keywords:
Topic-aware video summarization
Multimodal transformer
Video summarization dataset

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

B
beijing institute of technology
Scholars:
5.4W
Papers: 4.0W
Citations: 63