arrow
Return

Guilt-Free Data Reuse

delete2017-03-24
delete9
delete
OA
AI
C
Cynthia Dwork *
V
Vitaly Feldman
M
Moritz Hardt
T
Toniann Pitassi
O
Omer Reingold
A
Aaron Roth
DOI:10.1145/3051088delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
Existing approaches to ensuring the validity of inferences drawn from data assume a fixed procedure to be performed, selected before the data are examined. Yet the practice of data analysis is an intrinsically interactive and adaptive process: new analyses and hypotheses are proposed after seeing the results of previous ones, parameters are tuned on the basis of obtained results, and datasets are shared and reused. In this work, we initiate a principled study of how to guarantee the validity of statistical inference in adaptive data analysis. We demonstrate new approaches for addressing the challenges of adaptivity that are based on techniques developed in privacy-preserving data analysis. As an application of our techniques we give a simple and practical method for reusing a holdout (or testing) set to validate the accuracy of hypotheses produced adaptively by a learning algorithm operating on a training set.
Keywords:
STABILITY
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Communications of the ACM cover
Communications of the ACM
IF:
12.2
Papers:
1.2W
Citations:
3.7W

Organization

S
Stanford University
Scholars:
9.6W
Papers: 8.2W
Citations: 17.0W
U
university of pennsylvania
Scholars:
9.2W
Papers: 7.8W
Citations: 153
I
international business machines (ibm)
Scholars:
5.7K
Papers: 4.5K
Citations: 4
M
Microsoft
Scholars:
3.0K
Papers: 2.7K
Citations: 7
G
Google Incorporated
Scholars:
3.5K
Papers: 1.8K
Citations: 8
U
university of toronto
Scholars:
14.7W
Papers: 12.0W
Citations: 165
researcher View more organizations