arrow
返回

A multi-layer memory sharing network for video captioning

delete2023-04-01
delete11
PRE
AI
T
Tian-Zi Niu
S
Shan-Shan Dong
Z
Zhen-Duo Chen
X
Xin Luo
Z
Zi Huang
S
Shanqing Guo
X
Xin-Shun Xu *
DOI:10.1016/j.patcog.2022.109202delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Over the past several years, video captioning has received much attention in computer vision and ma-chine learning communities. Many models utilize an RNN-based decoder to generate sentences describing the content of a video. They have achieved much progress; however, few methods adopt a decoder with more than three layers because an RNN-based model with more layers may become hard to train, time-consuming or even deteriorate at a certain depth. To address the limitation, we propose a Multi-layer memory sharing Network, MesNet for short, which allows more layers to be stacked without compro-mising performance. In MesNet, we construct a novel memory sharing structure to strengthen the con-nections between layers and make the model easier to train. More specifically, we design an Enhanced Gated Recurrent Unit (En-GRU) and stack it to construct a deeper network. Unlike traditional RNN-based multi-layer networks, the memory states of all layers in MesNet are cross-used at each iteration to mimic the brain's complex connections. Extensive experiments on MSVD and MSR-VTT demonstrate that our method performs well and outperforms some state-of-the-art methods significantly. Our code is available at https://github.com/nbbb/MesNet .(c) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Video captioning
Multi -layer network
Memory sharing
Enhanced gated recurrent unit

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

S
shandong university
学者数:
9.4W
论文数: 6.4W
被引数: 94
U
University of Queensland
学者数:
5.0W
论文数: 5.1W
被引数: 9.2W
引用论文

引用论文

Human-Centric Image Captioning
err2022-06-01
err14
PREAI
errYang, Zuopeng; Wang, Pengbo; Chu, Tianshu; Yang, Jie
err分享
err收藏
Recent advances in convolutional neural networks卷积神经网络的最新进展
err2018-05-01
err3.8K
errOAAI
errGu, Jiuxiang; Wang, Zhenhua; Kuen, Jason; Ma, Lianyang; Shahroudy, Amir; Shuai, Bing; Liu, Ting; Wang, Xingxing; Wang, Gang; Cai, Jianfei; Chen, Tsuhan
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
SARS due to COVID-19: Predictors of death and profile of adult patients in the state of Rio de Janeiro, 2020
err2022-11-10
err0
errOAAI
errTatiana de Araujo Eleuterio; Marcella Cini Oliveira; Mariana dos Santos Velasco; Rachel de Almeida Menezes; Regina Bontorim Gomes; Marlos Melo Martins; Carlos Eduardo Raymundo; Roberto de Andrade Medronho
err分享
err收藏
Enhancing the alignment between target words and corresponding frames for video captioning
err2021-03-01
err41
PREAI
errTu, Yunbin; Zhou, Chang; Guo, Junjun; Gao, Shengxiang; Yu, Zhengtao
err分享
err收藏
学者 查看更多内容