arrow
Return

U-Transformer-based multi-levels refinement for weakly supervised action segmentation

delete2024-05-01
delete1
PRE
AI
肖克 cover
肖克 (Ke Xiao)
X
Xin Miao *
郭文忠 (Wenzhong Guo)
DOI:10.1016/j.patcog.2023.110199delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Action segmentation is a research hotspot in human action analysis, which aims to split videos into segments of different actions. Recent algorithms have achieved great success in modeling based on temporal convolution, but these methods weight local or global timing information through additional modules, ignoring the existing long-term and short-term information connections between actions. This paper proposes a U-Transformer structure based on multi-level refinement, introduces neighborhood attention to learn the neighborhood information of adjacent frames, and aggregates video frame features to effectively process long-term sequence information. Then a loss optimization strategy is proposed to smooth the original classification effect and generate a more accurate calibration sequence by introducing a pairing similarity optimization method based on deep feature learning. In addition, we propose a timestamp supervised training method to generate complete information for actions based on pseudo-label predictions for action boundary predictions. Experiments on three challenging action segmentation datasets, 50Salads, GTEA, and Breakfast, show that our model performs state-of-the-art models, and our weakly supervised model also performs comparably to fully supervised performance.
Keywords:
Action segmentation
U-Transformer
Timestamp supervision
Multi-stages refinement

Journal

Pattern Recognition cover
Pattern Recognition
IF:
7.6
Papers:
1.3W
Citations:
4.5W

Organization

F
fuzhou university
Scholars:
3.3W
Papers: 2.1W
Citations: 31