arrow
Return

SEMI-SUPERVISED U-STATISTICS

delete2025-12-01
delete0
PRE
AI
I
Ilmun Kim *
L
L A Wasserman
S
Sivaraman Balakrishnan
M
Matey Neykov
DOI:10.1214/25-AOS2550delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Semi-supervised datasets are ubiquitous across diverse domains where obtaining fully labeled data is costly or time-consuming. The prevalence of such datasets has consistently driven the demand for new tools and methods that exploit the potential of unlabeled data. Responding to this demand, we introduce semi-supervised U-statistics enhanced by the abundance of unlabeled data, and investigate their statistical properties. We show that the proposed approach is asymptotically Normal and exhibits notable efficiency gains over classical U-statistics by effectively integrating various powerful prediction tools into the framework. To understand the fundamental difficulty of the problem, we derive minimax lower bounds in semi-supervised settings and showcase that our procedure is semi-parametrically efficient under regularity conditions. Moreover, tailored to bivariate kernels, we propose a refined approach that outperforms the classical U-statistic across all degeneracy regimes, and demonstrate its optimality properties. Simulation studies are conducted to corroborate our findings and to further demonstrate our framework.
Keywords:
minimax risk
semi-supervised inference
U-statistics
van Trees inequal-ity
Adaptivity

Journal

Annals of Statistics cover
Annals of Statistics
IF:
3.7
Papers:
2.8K
Citations:
2.9W

Organization

N
northwestern university
Scholars:
4.5K
Papers: 1.8K
Citations: 1
C
carnegie mellon university
Scholars:
1.9K
Papers: 952
Citations: 0