arrow
返回

Deep multimodal embedding for video captioning

delete2019-07-24
delete9
PRE
AI
J
Jin Young Lee *
DOI:10.1007/s11042-019-08011-3delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Automatically generating natural language descriptions from videos, which is simply called video captioning, is very challenging work in computer vision. Thanks to the success of image captioning, in recent years, there has been rapid progress in the video captioning. Unlike images, videos have a variety of modality information, such as frames, motion, audio, and so on. However, since each modality has different characteristic, how they are embedded in a multimodal video captioning network is very important. This paper proposes a deep multimodal embedding network based on analysis of the multimodal features. The experimental results show that the captioning performance of the proposed network is very competitive in comparison with conventional networks.
Keyword:
Deep embedding
LSTM network
Multimodal features
Video captioning
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

Multimedia Tools and Applications 封面图
Multimedia Tools and Applications
IF:
3
论文数:
1.9W
被引数:
3.2W

机构

S
Sejong University
学者数:
8.3K
论文数: 1.1W
被引数: 1.5W
引用论文

引用论文

Beteiligung des peripheren Nervensystems bei Morbus Crohn
err1999-12-03
err0
PREAI
errB. Moormann; H. Herath; O. Mann; A. Ferbert
err分享
err收藏
Structural phase transitions in aluminium above 320 GPa
err2018-09-28
err0
errOAAI
errGuillaume Fiquet; Chandrabhas Narayana; Christophe Bellin; Abhay Shukla; Imène Estève; Art L. Ruoff; Gaston Garbarino; Mohamed Mezouar
err分享
err收藏
学者 查看更多内容