arrow
Return

RadlER: Deduplicated Sampling On-Demand

delete2025-08-01
delete0
PRE
AI
L
Luca Zecchini *
Z
Ziawasch Abedjan
V
Vasilis Efthymiou
G
Giovanni Simonini
DOI:10.14778/3750601.3750661delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Data practitioners often need to sample their datasets to producerepresentative subsets for their downstream tasks. Unfortunately,real-world datasets frequently contain duplicates, whose presencebiases sampling and impacts the quality of the produced subsets,hence the outcome of downstream tasks. While deduplication istherefore fundamental, performing it on the entire dataset to runsampling on its cleaned version might be prohibitively expensive interms of time and resources. Thus, we recently introducedRadlER,a solution to performdeduplicated sampling on-demand, i.e., toproduce a clean sample of a dirty dataset incrementally, accordingto a target distribution of some subpopulations, by focusing thecleaning effort only on entities required to appear in the sample.In this demonstration, we interactively show howRadlERcansupport practitioners in their data science pipelines, allowing themto save a relevant amount of time and resources

Journal

P
Proceedings of the VLDB Endowment
IF:
3.3
Papers:
556
Citations:
1.2W

Organization

U
universita di modena e reggio emilia
Scholars:
1.6W
Papers: 1.2W
Citations: 12
T
Technical University of Berlin
Scholars:
1.3W
Papers: 1.1W
Citations: 18