返回
State-aware video procedural captioning
DOI:10.1007/s11042-023-14774-7.png)
摘要
En 中文
Video procedural captioning (VPC), which generates procedural text from instructional videos, is an essential task for scene understanding and real-world applications. The main challenge of VPC is to describe how to manipulate materials accurately. This paper focuses on this challenge by designing a new VPC task, generating a procedural text from the clip sequence of an instructional video and material set. In this task, the state of materials is sequentially changed by manipulations, yielding their state-aware visual representations (e.g., eggs are transformed into cracked, stirred, then fried forms). The essential difficulty is to convert such visual representations into textual representations; that is, a model should track the material states after manipulations to better associate the cross-modal relations. To achieve this, we propose a novel VPC method, which modifies an existing textual simulator for tracking material states as a visual simulator and incorporates it into a video captioning model. Our experimental results show the effectiveness of the proposed method, which outperforms state-of-the-art video captioning models. We further analyze the learned embedding of materials to demonstrate that the simulators capture their state transition.
Keyword:
Instructional video
Procedural text
Simulator
期刊
IF:
3
论文数:
2.0W
被引数:
3.2W
机构
引用论文
TUMOR-NECROSIS-FACTOR AND DNA TOPOISOMERASE-II INHIBITORS IN HUMAN OVARIAN-CANCER - POTENTIAL ROLE IN CHEMOTHERAPY肿瘤坏死因子和DNA拓扑异构酶-II抑制剂在人类卵巢癌中的潜在作用:化疗中的应用
Gamma-delta T-cell Lymphoma with CNS Involvement Presenting with Proptosis: A Case Study Workup, Treatment and Prognosis
Orbit
IF0

