arrow
Return

TSRN: two-stage refinement network for temporal action segmentation

delete2023-05-15
delete10
PRE
AI
X
Xiaoyan Tian
Y
Ye Jin *
DOI:10.1007/s10044-023-01166-8delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In high-level video semantic understanding, continuous action segmentation is a challenging task aimed at segmenting an untrimmed video and labeling each segment with predefined labels over time. However, the accuracy of segment predictions is limited by confusing information in video sequences, such as ambiguous frames during action boundaries or over-segmentation errors due to the lack of semantic relations. In this work, we present a two-stage refinement network (TSRN) to improve temporal action segmentation. We first capture global relations over an entire video sequence using a multi-head self-attention mechanism in the novel transformer temporal convolutional network and model temporal relations in each action segment. Then, we introduce a dual-attention spatial pyramid pooling network to fuse features from macro-scale and microscale perspectives, providing more accurate classification results from the initial prediction. In addition, a joint loss function mitigates over-segmentation. Compared with state-of-the-art methods, the proposed TSRN substantially improves temporal action segmentation on three challenging datasets (i.e., 50Salads, Georgia Tech Egocentric Activities, and Breakfast).
Keywords:
Temporal action segmentation
Video semantic understanding
Refinement network
Self-attention
Over-segmentation

Journal

Pattern Analysis and Applications cover
Pattern Analysis and Applications
IF:
2
Papers:
1.9K
Citations:
1.9K

Organization

H
harbin institute of technology
Scholars:
8.0W
Papers: 6.6W
Citations: 66