Return
Less is More: Unlabeled Data Selection for Efficient Tabular Self-supervised Learning
DOI:10.1007/s10994-026-07133-8.png)
Abstract
En 中文
Self-supervised learning (SSL) addresses label scarcity by leveraging large amounts of unlabeled data to learn transferable representations. However, pretraining on very large unlabeled datasets can be computationally expensive and may include noisy or unrepresentative samples that degrade learning. In this work, we explore whether selecting a subset of the available unlabeled examples for the pretext task can reduce computational cost while maintaining strong performance of SSL methods for tabular data. In particular, we investigate whether uncertainty-, diversity-, and transport-based criteria can guide this selection and improve representation quality. To this end, we conduct large-scale experiments on 25 tabular benchmark datasets using four SSL models, across varying amounts of labeled data under both biased and unbiased label selection, and multiple strategies for sampling unlabeled data. To better understand when and how unlabeled data subsampling is effective, we relate dataset characteristics to experimentally observed performance gains through a meta-analysis. Our results show that reducing the unlabeled data pool yields substantial computational savings while maintaining, and often even improving, downstream performance.
Keywords:
Self-supervised learning
Tabular data
Unlabeled data selection
Active data selection
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
2.9
Papers:
2.7K
Citations:
3.4W
Organization
Cited Papers
no more

