arrow
返回

Repeat and learn: Self-supervised visual representations learning by Scene Localization

delete2024-12-01
delete0
PRE
AI
H
Hussein Altabrawee
M
Mohd Halim Mohd Noor *
DOI:10.1016/j.patcog.2024.110804delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Large labeled datasets are crucial for video understanding progress. However, the labeling process is timeconsuming, expensive, and tiresome. To overcome this impediment, various pretexts use the temporal coherence in videos to learn visual representations in a self-supervised manner. However, these pretexts (order verification and sequence sorting) struggle when encountering cyclic actions due to the label ambiguity problem. To overcome these limitations, we present a novel temporal pretext task to address self-supervised learning of visual representations from unlabeled videos. Repeated Scene Localization (RSL) is a multi-class classification pretext that involves changing the temporal order of the frames in a video by repeating a scene. Then, the network is trained to identify the modified video, localize the location of the repeated scene, and identify the unmodified original videos that do not have repeated scenes. We evaluated the proposed pretext on two benchmark datasets, UCF-101 and HMDB-51. The experimental results show that the proposed pretext achieves state-of-the-art results in action recognition and video retrieval tasks. In action recognition, our S3D model achieves 88.15% and 56.86% on UCF-101 and HMDB-51, respectively. It outperforms the current state-of-the-art by 1.05% and 3.26%. Our R(2+1)D-Adjacent model achieves 83.52% and 54.50% on UCF-101 and HMDB-51, respectively. It outperforms the single pretext tasks by 8.7% and 13.9%. In video retrieval, our R(2+1)D-Offset model outperforms the single pretext tasks by 4.68% and 1.1% Top 1 accuracies on UCF-101 and HMDB-51, respectively. The source code and the trained models are publicly available at https://github.com/Hussein-A-Hassan/RSL-Pretext.
Keyword:
Visual representations learning
Action recognition
Self-supervised learning

期刊

Pattern Recognition 封面图
Pattern Recognition
IF:
7.6
论文数:
1.3W
被引数:
4.5W

机构

U
Universiti Sains Malaysia
学者数:
1.5W
论文数: 1.3W
被引数: 131
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
err分享
err收藏
err分享
err收藏
Temporal segment dropout for human action video recognition用于人体动作视频识别的时间片段丢失
err2024-02-01
err4
PREAI
errZhang, Yu; Chen, Zhengjie; Xu, Tianyu; Zhao, Junjie; Mi, Siya; Geng, Xin; Zhang, Min-Ling
err分享
err收藏
学者 查看更多内容