arrow
返回

Learning deep spatiotemporal features for video captioning

delete2018-12-01
delete11
PRE
AI
E
Eleftherios Daskalakis
M
Maria Tzelepi *
A
Anastasios Tefas
DOI:10.1016/j.patrec.2018.09.022delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this paper, we propose a novel automatic video captioning system which translates videos to sentences, utilizing a deep neural network that is composed of three building parts of convolutional and recurrent structure. That is, the first subnetwork operates as feature extractor of single frames. The second subnetwork is a three-stream network, capable of capturing spatial semantic information in the first stream, temporal semantic information in the second stream, and global video concept information in the third stream. The third subnetwork generates relevant textual captions using as input the spatiotemporal features of the second subnetwork. The experimental validation indicates the effectiveness of the proposed model, achieving superior performance over competitive methods. (C) 2018 Elsevier B.V. All rights reserved.
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Pattern Recognition Letters 封面图
Pattern Recognition Letters
IF:
3.3
论文数:
7.9K
被引数:
1.6W

机构

A
aristotle university of thessaloniki
学者数:
2.6W
论文数: 2.0W
被引数: 19
引用论文

引用论文

ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
A region-based image caption generator with refined descriptions
err2018-01-01
err76
errOAAI
errKinghorn, Philip; Zhang, Li; Shao, Ling
err分享
err收藏