arrow
返回

Self-Supervised Video Representation Learning by Serial Restoration With Elastic Complexity

delete2024-01-01
delete2
PRE
AI
Z
Ziyu Chen
H
Hanli Wang *
C
Chang Wen Chen
DOI:10.1109/TMM.2023.3293727delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Self-supervised video representation learning leaves out heavy manual annotation by automatically excavating supervisory signals. Although contrastive learning based approaches exhibit superior performances, pretext task based approaches still deserve further study. This is because the pretext tasks exploit the nature of data and encourage feature extractors to learn spatiotemporal logic by discovering dependencies among video clips or cubes, without manual engineering on data augmentations or manual construction of contrastive pairs. To utilize chronological property more effectively and efficiently, this work proposes a novel pretext task, named serial restoration of shuffled clips (SRSC), disentangled by an elaborately designed task network composed of an order-aware encoder and a serial restoration decoder. In contrast to other order based pretext tasks that formulate clip order recognition as a one-step classification problem, the proposed SRSC task restores shuffled clips into the right order in multiple steps. Owing to the excellent elasticity of SRSC, a novel taxonomy of curriculum learning is further proposed to equip SRSC with different pre-training strategies. According to the factors that affect the complexity of solving the SRSC task, the proposed curriculum learning strategies can be categorized into task based, model based and data based. Extensive experiments are conducted on the subdivided strategies to explore their effectiveness and noteworthy laws. Compared with existing approaches, this work demonstrates that the proposed approach achieves state-of-the-art performances in pretext task based self-supervised video representation learning and a majority of the proposed strategies further boost the performance of downstream tasks. For the first time, the features pre-trained by the pretext tasks are applied to video captioning by feature-level early fusion, and enhance the input of existing approaches as a lightweight plugin.
Keyword:
Self-supervised learning
pretext task
video representation learning
curriculum learning
video captioning
action recognition
nearest neighbor retrieval

期刊

IEEE Transactions on Multimedia 封面图
IEEE Transactions on Multimedia
IF:
9.7
论文数:
4.5K
被引数:
2.4W

机构

H
hong kong polytechnic university
学者数:
3.0W
论文数: 4.1W
被引数: 921
T
tongji university
学者数:
7.9W
论文数: 6.0W
被引数: 98
引用论文

引用论文

Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video Applications
err2021-01-01
err9
PREAI
errXiao, Jun; Li, Lin; Xu, Dejing; Long, Chengjiang; Shao, Jian; Zhang, Shifeng; Pu, Shiliang; Zhuang, Yueting
err分享
err收藏
A segmentation approach for the reproducible extraction and quantification of knickpoints from river long profiles
err2019-02-18
err0
errOAAI
errBoris Gailleton; Simon M. Mudd; Fiona J. Clubb; Daniel Peifer; Martin D. Hurst
err分享
err收藏
Essential considerations for exploring visual working memory storage in the human brain
err2021-05-05
err0
PREAI
errPolina Iamshchinina; Thomas B. Christophel; Surya Gayet; Rosanne L. Rademaker
err分享
err收藏
err分享
err收藏
Self-Supervised Graph Convolutional Network for Multi-View Clustering
err2022-01-01
err73
PREAI
errXia, Wei; Wang, Qianqian; Gao, Quanxue; Zhang, Xiangdong; Gao, Xinbo
err分享
err收藏
err分享
err收藏
Diagnostic ability of confocal near-infrared reflectance fundus imaging to detect retrograde microcystic maculopathy from chiasm compression. A comparative study with OCT findings
err2021-06-24
err0
errOAAI
errMário L. R. Monteiro; Rafael M. Sousa; Rafael B. Araújo; Daniel Ferraz; Mohammad A. Sadiq; Leandro C. Zacharias; Rony C. Preti; Leonardo P. Cunha; Quan D. Nguyen
err分享
err收藏
A fast mesh deformation method for marine propeller flow
err2019-02-11
err0
PREAI
errJize Zhong; Zhiqiang Xie; Chunxu Wang; Du Shen; Hao Wang
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
学者 查看更多内容