返回
Classifying data from protected statistical datasets
DOI:10.1016/j.cose.2010.05.005.png)
摘要
En 中文
Statistical Disclosure Control (SDC) is an active research area in the recent years. The goal is to transform an original dataset X into a protected one X', such that X' does not reveal any relation between confidential and (quasi-)identifier attributes and such that X' can be used to compute reliable statistical information about X. Many specific protection methods have been proposed and analyzed, with respect to the levels of privacy and utility that they offer. However, when measuring utility, only differences between the statistical values of X and X' are considered. This would indicate that datasets protected by SDC methods can be used only for statistical purposes. We show in this paper that this is not the case, because a protected dataset X' can be used to construct good classifiers for future data. To do so, we describe an extensive set of experiments that we have run with different SDC protection methods and different (real) datasets. In general, the resulting classifiers are very good, which is good news for both the SDC and the Privacy-preserving Data Mining communities. In particular, our results question the necessity of some specific protection methods that have appeared in the privacy-preserving data mining (PPDM) literature with the clear goal of providing good classification. (C) 2010 Elsevier Ltd. All rights reserved.
Keyword:
Statistical disclosure control
Classification methods
Disclosure risk
Information loss
WEKA experiments
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
C
IF:
5.4
论文数:
4.6K
被引数:
1.4W
机构
引用论文
Electrical Conductivity of Reproductive Tissue for Detection of Estrus in Dairy Cows用于检测奶牛发情的繁殖组织电导率
Medical prescribing and antibiotic resistance: A game-theoretic analysis of a potentially catastrophic social dilemma
PLOS ONE
IF0

