Return
RadlER: Deduplicated Sampling On-Demand
DOI:10.14778/3750601.3750661.png)
Abstract
En 中文
Data practitioners often need to sample their datasets to producerepresentative subsets for their downstream tasks. Unfortunately,real-world datasets frequently contain duplicates, whose presencebiases sampling and impacts the quality of the produced subsets,hence the outcome of downstream tasks. While deduplication istherefore fundamental, performing it on the entire dataset to runsampling on its cleaned version might be prohibitively expensive interms of time and resources. Thus, we recently introducedRadlER,a solution to performdeduplicated sampling on-demand, i.e., toproduce a clean sample of a dirty dataset incrementally, accordingto a target distribution of some subpopulations, by focusing thecleaning effort only on entities required to appear in the sample.In this demonstration, we interactively show howRadlERcansupport practitioners in their data science pipelines, allowing themto save a relevant amount of time and resources
Journal
P
IF:
3.3
Papers:
556
Citations:
1.2W

