arrow
Return

A Dataset Auditing Method for Collaboratively Trained Machine Learning Models

delete2023-07-01
delete4
PRE
AI
Y
Yangsibo Huang *
C
Chun-Yin Huang
X
Xiaoxiao Li *
K
Kai Li
DOI:10.1109/TMI.2022.3220706delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Dataset auditing for machine learning (ML) models is a method to evaluate if a given dataset is used in training a model. In a Federated Learning setting where multiple institutions collaboratively train a model with their decentralized private datasets, dataset auditing can facilitate the enforcement of regulations, which provide rules for preserving privacy, but also allow users to revoke authorizations and remove their data from collaboratively trained models. This paper first proposes a set of requirements for a practical dataset auditing method, and then present a novel dataset auditing method called Ensembled Membership Auditing (EMA). Its key idea is to leverage previously proposed Membership Inference Attack methods and to aggregate data-wise membership scores using statistic testing to audit a dataset for a ML model. We have experimentally evaluated the proposed approach with benchmark datasets, as well as 4 X-ray datasets (CBIS-DDSM, COVIDx, Child-XRay, and CXR-NIH) and 3 dermatology datasets (DERM7pt, HAM10000, and PAD-UFES-20). Our results show that EMA meet the requirements substantially better than the previous state-of-the-art method. Our code is at:https://github.com/Hazelsuko07/EMA.
Keywords:
Privacy
dataset auditing
medical image classification

Journal

IEEE Transactions on Medical Imaging cover
IEEE Transactions on Medical Imaging
IF:
9.8
Papers:
6.2K
Citations:
3.7W

Organization

P
Princeton University
Scholars:
2.1W
Papers: 2.3W
Citations: 5.1W
U
University of British Columbia
Scholars:
6.9W
Papers: 6.1W
Citations: 8.6W