Return
Sampling for computational efficiency when conducting analyses in big data
DOI:10.1093/aje/kwaf268.png)
Abstract
En 中文
A challenge to research in big data is the inherent computational intensity of analyses, particularly when using rigorous methods to address biases. We demonstrate the use of sampling methods in big data to estimate parameters using fewer resources. Our motivating question was whether lung cancer incidence differs by baseline HIV status, using a cohort of nearly 30 million Medicaid beneficiaries. We targeted three parameters (with listed estimator): incidence rate ratio (IRR, Poisson model), HR (Cox model), and risk ratio (RR, Kaplan–Meier). We controlled for confounders using inverse probability weighting. We ran analyses using the full sample and several sampling schemes: divide-and-recombine (10, 20, 50 samples), subcohort, and case-cohort. We compared point estimates, standard errors, computation time, and memory used. We observed 1113 incident lung cancer diagnoses among 180 980 beneficiaries with HIV and 33 106 diagnoses among 29 179 940 beneficiaries without HIV. Findings were similar across target parameters. The subcohort and case-cohort approaches had estimates closer to the full sample and were faster and less memory intensive than divide-and-recombine, especially when estimating the risk ratio. Including nonsampled cases in the case-cohort resulted in increases in computation time and memory relative to the subcohort approach.
Keywords:
Big data
Sampling methods
Computational efficiency
Bias correction
Parameter estimation
AI Summary
Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.
Journal
IF:
4.8
Papers:
9.9K
Citations:
3.7W

