arrow
Return

Multi-view and multi-augmentation for self-supervised visual representation learning

delete2023-12-16
delete3
PRE
AI
V
Van Nhiem Tran
C
Chi-En Huang
S
Shen-Hsuan Liu
M
Muhammad Saqlain Aslam
K
Kailin Yang
Y
Yung‐Hui Li *
J
Jia‐Ching Wang
DOI:10.1007/s10489-023-05163-6delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
In the real world, the appearance of identical objects depends on factors as varied as resolution, angle, illumination conditions, and viewing perspectives. This suggests that the data augmentation pipeline could benefit downstream tasks by exploring the overall data appearance in a self-supervised framework. Previous work on self-supervised learning that yields outstanding performance relies heavily on data augmentation such as cropping and color distortion. However, most methods use a static data augmentation pipeline, limiting the amount of feature exploration. To generate representations that encompass scale-invariant, explicit information about various semantic features and are invariant to nuisance factors such as relative object location, brightness, and color distortion, we propose the Multi-View, Multi-Augmentation (MVMA) framework. MVMA consists of multiple augmentation pipelines, with each pipeline comprising an assortment of augmentation policies. By refining the baseline self-supervised framework to investigate a broader range of image appearances through modified loss objective functions, MVMA enhances the exploration of image features through diverse data augmentation techniques. Transferring the resultant representation learning using convolutional networks (ConvNets) to downstream tasks yields significant improvements compared to the state-of-the-art DINO across a wide range of vision tasks and classification tasks: +4.1% and +8.8% top-1 on the ImageNet dataset with linear evaluation and k-NN classifier, respectively. Moreover, MVMA achieves a significant improvement of +5% AP(50) and +7% AP(50)(m) on COCO object detection and segmentation.
Keywords:
Multi-augmentation
SSL augmentation pipelines
Data augmentation policies
Nuisance factors
Scale-invariant representation learning
Metric learning

Journal

Applied Intelligence cover
Applied Intelligence
IF:
3.5
Papers:
7.6K
Citations:
1.7W

Organization

F
foxconn
Scholars:
97
Papers: 88
Citations: 0
N
National Central University
Scholars:
1.0W
Papers: 8.6K
Citations: 6.4K
Cited Papers

Cited Papers

Intra-Abdominal Splenosis Mimicking Metastatic Cancer
err2011-03-01
err0
PREAI
errNicholas J. Short; Teresa G. Hayes; Peeyush Bhargava
errShare
errSave
ImageNet Large Scale Visual Recognition Challenge
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
errShare
errSave
Locomotor deficits in the mutant mouse, Lurcher
err1987-04-01
err0
PREAI
errP.A. Fortier; A.M. Smith; S. Rossignol
errShare
errSave
Predictive Models for Cochlear Implantation in Elderly Candidates
err2005-12-01
err0
PREAI
errJanice Leung; Nae-Yuh Wang; Jennifer D. Yeagle; Jill Chinnici; Stephen Bowditch; Howard W. Francis; John K. Niparko
errShare
errSave
Structural study of lanthanides(III) in aqueous nitrate and chloride solutions by EXAFS
err1999-02-01
err0
PREAI
errT. Yaita; H. Narita; Sh. Suzuki; Sh. Tachimori; H. Motohashi; H. Shiwaku
errShare
errSave
ImageNet Classification with Deep Convolutional Neural Networks
err2017-05-24
err8.3W
errOAAI
errKrizhevsky, Alex; Sutskever, Ilya; Hinton, Geoffrey E.
errShare
errSave
Heuristic Attention Representation Learning for Self-Supervised Pretraining
errSENSORS
IF3.5
err2022-07-10
err4
errOAAI
errVan Nhiem Tran; Liu, Shen-Hsuan; Li, Yung-Hui; Wang, Jia-Ching
errShare
errSave
The role of natural killer T cells in a mouse model with spontaneous bile duct inflammation
err2017-02-20
err0
errOAAI
errElisabeth Schrumpf; Xiaojun Jiang; Sebastian Zeissig; Marion J. Pollheimer; Jarl Andreas Anmarkrud; Corey Tan; Mark A. Exley; Tom H. Karlsen; Richard S. Blumberg; Espen Melum
errShare
errSave
researcher View more