返回
Movie Description
DOI:10.1007/s11263-016-0987-1.png)
摘要
En 中文
Audio description (AD) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting data source for computer vision and computational linguistics. In this work we propose a novel dataset which contains transcribed ADs, which are temporally aligned to full length movies. In addition we also collected and aligned movie scripts used in prior work and compare the two sources of descriptions. We introduce the Large Scale Movie Description Challenge (LSMDC) which contains a parallel corpus of 128,118 sentences aligned to video clips from 200 movies (around 150 h of video in total). The goal of the challenge is to automatically generate descriptions for the movie clips. First we characterize the dataset by benchmarking different approaches for generating video descriptions. Comparing ADs to scripts, we find that ADs are more visual and describe precisely what is shown rather than what should happen according to the scripts created prior to movie production. Furthermore, we present and compare the results of several teams who participated in the challenges organized in the context of two workshops at ICCV 2015 and ECCV 2016.
Keyword:
Movie description
Video description
Video captioning
Video understanding
Movie description dataset
Movie description challenge
Long short-term
memory network
Audio description
LSMDC
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
9.3
论文数:
3.9K
被引数:
2.8W
机构
引用论文
The role of a critical left fronto-temporal network with its right-hemispheric homologue in syntactic learning based on word category information基于单词类别信息的关键左前-时间网络及其右半球同源物在句法学习中的作用

