arrow
Return

Less is More: Unlabeled Data Selection for Efficient Tabular Self-supervised Learning

delete2026-08-29
delete0
delete
OA
AI
S
Sintija Stevanoska *
C
Christian L. Camacho Villalón
S
Sašo Džeroski
K
Katharina Dost
DOI:10.1007/s10994-026-07133-8delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Self-supervised learning (SSL) addresses label scarcity by leveraging large amounts of unlabeled data to learn transferable representations. However, pretraining on very large unlabeled datasets can be computationally expensive and may include noisy or unrepresentative samples that degrade learning. In this work, we explore whether selecting a subset of the available unlabeled examples for the pretext task can reduce computational cost while maintaining strong performance of SSL methods for tabular data. In particular, we investigate whether uncertainty-, diversity-, and transport-based criteria can guide this selection and improve representation quality. To this end, we conduct large-scale experiments on 25 tabular benchmark datasets using four SSL models, across varying amounts of labeled data under both biased and unbiased label selection, and multiple strategies for sampling unlabeled data. To better understand when and how unlabeled data subsampling is effective, we relate dataset characteristics to experimentally observed performance gains through a meta-analysis. Our results show that reducing the unlabeled data pool yields substantial computational savings while maintaining, and often even improving, downstream performance.
Keywords:
Self-supervised learning
Tabular data
Unlabeled data selection
Active data selection
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Machine Learning cover
Machine Learning
IF:
2.9
Papers:
2.7K
Citations:
3.4W

Organization

J
jožef stefan institute
Scholars:
723
Papers: 318
Citations: 0
U
university of canterbury
Scholars:
1.2K
Papers: 636
Citations: 0
Cited Papers

Cited Papers

Random Forests
err
IF0
err2001-01-01
err0
PREAI
errLeo Breiman
errShare
errSave
Self-Supervised Representation Learning: Introduction, advances, and challenges
err2022-05-01
err174
errOAAI
errEricsson, Linus; Gouk, Henry; Loy, Chen Change; Hospedales, Timothy M.
errShare
errSave
Subset selection for domain adaptive pre-training of language model
err2025-03-19
err0
errOAAI
errHwang, Junha; Lee, Seungdong; Kim, Haneul; Jeong, Young-Seob
errShare
errSave
Hitting the target: stopping active learning at the cost-based optimum
err2022-10-14
err4
errOAAI
errPullar-Strecker, Zac; Dost, Katharina; Frank, Eibe; Wicker, Joerg
errShare
errSave
OpenML
err2014-06-16
err0
errOAAI
errJoaquin Vanschoren; Jan N. van Rijn; Bernd Bischl; Luis Torgo
errShare
errSave
errShare
errSave
no more