arrow
Return

Structure-Aware Procedural Text Generation From an Image Sequence

delete2021-01-01
delete8
delete
OA
AI
T
Taichi Nishimura *
A
Atsushi Hashimoto
Y
Yoshitaka Ushiku
H
Hirotaka Kameko
Y
Yoko Yamakata
S
Shinsuke Mori
DOI:10.1109/ACCESS.2020.3043452delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
It is an important activity for our society to create new value by combining materials. From daily cooking to manufacturing for industry, we often describe the way to do it as a procedural text. As pointed by some previous studies for natural language understanding, one important property of the procedural text is its dependency of the context, which is the merging operations of materials and can be represented by a graph or tree structure. This paper aims to investigate the impact of explicitly introducing such a structure on the vision and language task of procedural text generation from an image sequence. To this end, we propose (1) a new dataset, which extends a definition of a tree structure merging tree to a vision and language version and (2) a novel structure-aware procedural text generation model, which learns the context dependency efficiently. Experimental results show that the proposed method can boost the performance of traditional versatile methods.
Keywords:
Image sequences
Merging
Videos
Task analysis
Visualization
Oils
Annotations
Natural language processing
text generation
procedural text
vision and language
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

IEEE Access cover
IEEE Access
IF:
3.6
Papers:
9.8W
Citations:
29.4W

Organization

K
Kyoto University
Scholars:
5.1W
Papers: 4.6W
Citations: 6.1W
U
University of Tokyo
Scholars:
7.1W
Papers: 6.5W
Citations: 2.2K