arrow
Return

Self-Supervised Scene-Debiasing for Video Representation Learning via Background Patching

delete2023-01-01
delete11
PRE
AI
M
Maregu Assefa
W
Wei Jiang *
K
Kumie Gedamu
G
Getinet Yilma
DOI:10.1109/TMM.2022.3193559delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Self-supervised learning has considerably improved video representation learning by discovering supervisory signals automatically from unlabeled videos. However, due to the scene-biased nature of existing video datasets, the current methods are biased to the dominant scene context during action inference. Hence, this paper proposes Background Patching (BP), a scene-debiasing augmentation strategy to alleviate the model reliance on the video background in a self-supervised contrastive manner. The BP reduces the negative influence of the video background by mixing a randomly patched frame to the video background. BP randomly crops four frames from four different videos and patches them to construct a new frame for each video separately. The patched frame is mixed with all frames of the target video to produce a spatially distorted video sample. Then, we use existing self-supervised contrastive frameworks to pull representations of the distorted and original videos closer together. Moreover, BP mixes the semantic labels of patches with the target video's label, resulting in the regularization of the contrastive model to soften the decision boundaries in the embedding space. Therefore, the model is explicitly constrained to suppress the background influence by emphasizing more on the motion changes. The extensive experimental results show that our BP significantly improved the performance of various video understanding downstream tasks including action recognition, action detection, and video retrieval.
Keywords:
Action recognition
background patching
label smoothing
scene-debiasing
self-supervised learning
video representation

Journal

IEEE Transactions on Multimedia cover
IEEE Transactions on Multimedia
IF:
9.7
Papers:
4.5K
Citations:
2.4W

Organization

No organization information available
Cited Papers

Cited Papers

err
IF0
err
err0
PREAI
err
errShare
errSave
Time-dependent depolarization of aligned HD molecules
err2009-01-01
err0
errOAAI
errNate C.-M. Bartlett; Daniel J. Miller; Richard N. Zare; Andrew J. Alexander; Dimitris Sofikitis; T. Peter Rakitzis
errShare
errSave
Explore Video Clip Order With Self-Supervised and Curriculum Learning for Video Applications
err2021-01-01
err9
PREAI
errXiao, Jun; Li, Lin; Xu, Dejing; Long, Chengjiang; Shao, Jian; Zhang, Shifeng; Pu, Shiliang; Zhuang, Yueting
errShare
errSave
Spatio-Temporal Attention Networks for Action Recognition and Detection
err2020-11-01
err113
PREAI
errLi, Jun; Liu, Xianglong; Zhang, Wenxuan; Zhang, Mingyuan; Song, Jingkuan; Sebe, Nicu
errShare
errSave
Genomes of Two Chronological Isolates (Helicobacter pylori 2017 and 2018) of the West African Helicobacter pylori Strain 908 Obtained from a Single Patient
err2011-07-01
err0
errOAAI
errTiruvayipati Suma Avasthi; Singamaneni Haritha Devi; Todd D. Taylor; Narender Kumar; Ramani Baddam; Shinji Kondo; Yutaka Suzuki; Hervé Lamouliatte; Francis Mégraud; Niyaz Ahmed
errShare
errSave
researcher View more