返回
A multi-layer memory sharing network for video captioning
DOI:10.1016/j.patcog.2022.109202.png)
摘要
En 中文
Over the past several years, video captioning has received much attention in computer vision and ma-chine learning communities. Many models utilize an RNN-based decoder to generate sentences describing the content of a video. They have achieved much progress; however, few methods adopt a decoder with more than three layers because an RNN-based model with more layers may become hard to train, time-consuming or even deteriorate at a certain depth. To address the limitation, we propose a Multi-layer memory sharing Network, MesNet for short, which allows more layers to be stacked without compro-mising performance. In MesNet, we construct a novel memory sharing structure to strengthen the con-nections between layers and make the model easier to train. More specifically, we design an Enhanced Gated Recurrent Unit (En-GRU) and stack it to construct a deeper network. Unlike traditional RNN-based multi-layer networks, the memory states of all layers in MesNet are cross-used at each iteration to mimic the brain's complex connections. Extensive experiments on MSVD and MSR-VTT demonstrate that our method performs well and outperforms some state-of-the-art methods significantly. Our code is available at https://github.com/nbbb/MesNet .(c) 2022 Elsevier Ltd. All rights reserved.
Keyword:
Video captioning
Multi -layer network
Memory sharing
Enhanced gated recurrent unit
期刊
IF:
7.6
论文数:
1.3W
被引数:
4.5W
机构
引用论文
GR-RNN: Global-context residual recurrent neural networks for writer identification
PATTERN RECOGNITION
IF7.6
SARS due to COVID-19: Predictors of death and profile of adult patients in the state of Rio de Janeiro, 2020
PLOS ONE
IF0
Enhancing the alignment between target words and corresponding frames for video captioning
PATTERN RECOGNITION
IF7.6
Optimized Deep Stacked Long Short-Term Memory Network for Long-Term Load Forecasting用于长期负荷预测的优化深度堆叠长短期记忆网络
IEEE ACCESS
IF3.6

