arrow
返回

Multi-view and multi-augmentation for self-supervised visual representation learning

delete2023-12-16
delete3
PRE
AI
V
Van Nhiem Tran
C
Chi-En Huang
S
Shen-Hsuan Liu
M
Muhammad Saqlain Aslam
K
Kailin Yang
Y
Yung‐Hui Li *
J
Jia‐Ching Wang
DOI:10.1007/s10489-023-05163-6delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In the real world, the appearance of identical objects depends on factors as varied as resolution, angle, illumination conditions, and viewing perspectives. This suggests that the data augmentation pipeline could benefit downstream tasks by exploring the overall data appearance in a self-supervised framework. Previous work on self-supervised learning that yields outstanding performance relies heavily on data augmentation such as cropping and color distortion. However, most methods use a static data augmentation pipeline, limiting the amount of feature exploration. To generate representations that encompass scale-invariant, explicit information about various semantic features and are invariant to nuisance factors such as relative object location, brightness, and color distortion, we propose the Multi-View, Multi-Augmentation (MVMA) framework. MVMA consists of multiple augmentation pipelines, with each pipeline comprising an assortment of augmentation policies. By refining the baseline self-supervised framework to investigate a broader range of image appearances through modified loss objective functions, MVMA enhances the exploration of image features through diverse data augmentation techniques. Transferring the resultant representation learning using convolutional networks (ConvNets) to downstream tasks yields significant improvements compared to the state-of-the-art DINO across a wide range of vision tasks and classification tasks: +4.1% and +8.8% top-1 on the ImageNet dataset with linear evaluation and k-NN classifier, respectively. Moreover, MVMA achieves a significant improvement of +5% AP(50) and +7% AP(50)(m) on COCO object detection and segmentation.
Keyword:
Multi-augmentation
SSL augmentation pipelines
Data augmentation policies
Nuisance factors
Scale-invariant representation learning
Metric learning

期刊

Applied Intelligence 封面图
Applied Intelligence
IF:
3.5
论文数:
7.6K
被引数:
1.7W

机构

F
foxconn
学者数:
97
论文数: 88
被引数: 0
N
National Central University
学者数:
1.0W
论文数: 8.6K
被引数: 6.4K
引用论文

引用论文

Intra-Abdominal Splenosis Mimicking Metastatic Cancer
err2011-03-01
err0
PREAI
errNicholas J. Short; Teresa G. Hayes; Peeyush Bhargava
err分享
err收藏
ImageNet Large Scale Visual Recognition ChallengeImageNet大规模视觉识别挑战
err2015-04-11
err2.7W
PREAI
errRussakovsky, Olga; Deng, Jia; Su, Hao; Krause, Jonathan; Satheesh, Sanjeev; Ma, Sean; Huang, Zhiheng; Karpathy, Andrej; Khosla, Aditya; Bernstein, Michael; Berg, Alexander C.; Fei-Fei, Li
err分享
err收藏
Locomotor deficits in the mutant mouse, Lurcher
err1987-04-01
err0
PREAI
errP.A. Fortier; A.M. Smith; S. Rossignol
err分享
err收藏
Predictive Models for Cochlear Implantation in Elderly Candidates
err2005-12-01
err0
PREAI
errJanice Leung; Nae-Yuh Wang; Jennifer D. Yeagle; Jill Chinnici; Stephen Bowditch; Howard W. Francis; John K. Niparko
err分享
err收藏
err分享
err收藏
Heuristic Attention Representation Learning for Self-Supervised Pretraining
errSENSORS
IF3.5
err2022-07-10
err4
errOAAI
errVan Nhiem Tran; Liu, Shen-Hsuan; Li, Yung-Hui; Wang, Jia-Ching
err分享
err收藏
The role of natural killer T cells in a mouse model with spontaneous bile duct inflammation
err2017-02-20
err0
errOAAI
errElisabeth Schrumpf; Xiaojun Jiang; Sebastian Zeissig; Marion J. Pollheimer; Jarl Andreas Anmarkrud; Corey Tan; Mark A. Exley; Tom H. Karlsen; Richard S. Blumberg; Espen Melum
err分享
err收藏
学者 查看更多内容