arrow
返回

Parallel Pathway Dense Video Captioning With Deformable Transformer

delete2022-01-01
delete6
delete
OA
AI
W
Wangyu Choi
J
Jiasi Chen
J
Jongwon Yoon *
DOI:10.1109/ACCESS.2022.3228821delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Dense video captioning is a very challenging task because it requires a high-level understanding of the video story, as well as pinpointing details such as objects and motions for a consistent and fluent description of the video. Many existing solutions divide this problem into two sub-tasks, event detection and captioning, and solve them sequentially ( localize-then-describe or reverse). Consequently, the final outcome is highly dependent on the performance of the preceding modules. In this paper, we decompose this sequential approach by proposing a parallel pathway dense video captioning framework that localizes and describes events simultaneously without any bottlenecks. We introduce a representation organization network at the branching point of the parallel pathway to organize the encoded video feature by considering the entire storyline. Then, an event localizer focuses to localize events without any event proposal generation network, a sentence generator describes events while considering the fluency and coherency of sentences. Our method has several advantages over existing work: (i) the final output does not depend on the output of the preceding modules, (ii) it improves existing parallel decoding methods by relieving the bottleneck of information. We evaluate the performance of PPVC on large-scale benchmark datasets, the ActivityNet Captions, and YouCook2. PPVC not only outperforms existing algorithms on the majority of metrics but also improves on both datasets by 5.4% and 4.9% compared to the state-of-the-art parallel decoding method.
Keyword:
Machine learning
deep learning
video and language
video captioning

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

H
hanyang university
学者数:
2.9W
论文数: 2.7W
被引数: 36
University of California System 封面图
University of California System
学者数:
37.6W
论文数: 33.8W
被引数: 6.6K
引用论文

引用论文

err分享
err收藏
err分享
err收藏
Understanding Objects in Video: Object-Oriented Video Captioning via Structured Trajectory and Adversarial Learning
err2020-01-01
err8
errOAAI
errZhu, Fangyi; Hwang, Jenq-Neng; Ma, Zhanyu; Chen, Guang; Guo, Jun
err分享
err收藏
Structural phase transitions in aluminium above 320 GPa
err2018-09-28
err0
errOAAI
errGuillaume Fiquet; Chandrabhas Narayana; Christophe Bellin; Abhay Shukla; Imène Estève; Art L. Ruoff; Gaston Garbarino; Mohamed Mezouar
err分享
err收藏
Climate policy: Steps to China's carbon peak气候政策: 迈向中国碳峰值的步骤
err2015-06-17
err0
errOAAI
errZhu Liu; Dabo Guan; Scott Moore; Henry Lee; Jun Su; Qiang Zhang
err分享
err收藏
err分享
err收藏
学者 查看更多内容