返回
Attention based video captioning framework for Hindi
DOI:10.1007/s00530-021-00816-3.png)
摘要
En 中文
In recent times, active research is going on for bridging the gap between computer vision and natural language. In this paper, we attempt to address the problem of Hindi video captioning. In a linguistically diverse country like India, it is important to provide a means which can help in understanding the visual entities in native languages. In this work, we employ a hybrid attention mechanism by extending the soft temporal attention mechanism with a semantic attention to make the system able to decide when to focus on visual context vector and semantic input. The visual context vector of the input video is extracted using 3D convolutional neural network (3D CNN) and a Long Short-Term Memory (LSTM) recurrent network with attention module is used for decoding the encoded context vector. We experimented on a dataset built in-house for Hindi video captioning by translating MSR-VTT\ dataset followed by post-editing. Our system achieves 0.369 CIDEr score and 0.393 METEOR score and outperformed other baseline models including RMN (Reasoning Module Networks)-based model.
Keyword:
Deep learning
Hindi captions
LSTM
Hindi video captioning
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.1
论文数:
2.8K
被引数:
2.7K
机构
引用论文
A pilot study on EORTC or PERCIST for the prediction of progression-free survival with nivolumab therapy in advanced or metastatic gastric cancers: A STROBE-compliant article一项关于EORTC或PERCIST在预测接受纳武利尤单抗治疗的晚期或转移性胃癌患者无进展生存期方面的试点研究:一篇符合STROBE指南的文章
A novel and efficient xanthenic dye–organometallic ion‐pair complex for photoinitiating polymerization一种用于光引发聚合的新型高效的黄原胶染料-有机金属离子对配合物
The role of a critical left fronto-temporal network with its right-hemispheric homologue in syntactic learning based on word category information基于单词类别信息的关键左前-时间网络及其右半球同源物在句法学习中的作用

