arrow
返回

Multistage temporal convolution transformer for action segmentation

delete2022-12-01
delete15
PRE
AI
N
Nicolas Aziere *
S
Siniša Todorović
DOI:10.1016/j.imavis.2022.104567delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper addresses fully supervised action segmentation. Transformers have been shown to have large model capacity and powerful sequence modeling abilities, and hence seem quite suitable for capturing action grammar in videos. However, their performance in video understanding still lags behind that of temporal convolutional networks, or ConvNets for short. We hypothesize that this is because: (i) ConvNets tend to generalize better than Transformers, and (ii) Transformer's large model capacity requires significantly larger training datasets than existing action segmentation benchmarks. We specify a new hybrid model, TCTr, that combines the strengths from both frameworks. TCTr seamlessly unifies depth-wise convolution and self-attention in a principled manner. Also, TCTr addresses the Transformer's quadratic computational and memory complexity in the sequence length by learning how to adaptively estimate attention from local temporal neighborhoods, instead of all frames. Our experiments show that TCTr significantly outperforms the state of the art on the Breakfast, GTEA, and 50Salads datasets.(c) 2022 Elsevier B.V. All rights reserved.
Keyword:
Action segmentation
Video understanding
Full supervision
Transformer network
Hybrid models
CNNs

期刊

Image and Vision Computing 封面图
Image and Vision Computing
IF:
4.2
论文数:
4.1K
被引数:
6.7K

机构

O
Oregon State University
学者数:
1.7W
论文数: 1.5W
被引数: 2.4W
引用论文

引用论文

err分享
err收藏