返回
Deep multimodal embedding for video captioning
DOI:10.1007/s11042-019-08011-3.png)
摘要
En 中文
Automatically generating natural language descriptions from videos, which is simply called video captioning, is very challenging work in computer vision. Thanks to the success of image captioning, in recent years, there has been rapid progress in the video captioning. Unlike images, videos have a variety of modality information, such as frames, motion, audio, and so on. However, since each modality has different characteristic, how they are embedded in a multimodal video captioning network is very important. This paper proposes a deep multimodal embedding network based on analysis of the multimodal features. The experimental results show that the captioning performance of the proposed network is very competitive in comparison with conventional networks.
Keyword:
Deep embedding
LSTM network
Multimodal features
Video captioning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3
论文数:
1.9W
被引数:
3.2W
机构
引用论文
Advancement of reproductive activity, seasonal reduction in prolactin secretion and seasonal pelage changes in pubertal red deer hinds (Cervus elaphus) subjected to artificially shortened daily photoperiod or daily melatonin treatments
Reproduction
IF0
A novel and efficient xanthenic dye–organometallic ion‐pair complex for photoinitiating polymerization一种用于光引发聚合的新型高效的黄原胶染料-有机金属离子对配合物
Development of Criteria for Lamina Emergent Mechanism Flexures With Specific Application to Metals金属板材突起机制弯曲发展的标准制定
The role of a critical left fronto-temporal network with its right-hemispheric homologue in syntactic learning based on word category information基于单词类别信息的关键左前-时间网络及其右半球同源物在句法学习中的作用

