arrow
返回

Bayesian Variable Selection in Linear Regression in One Pass for Large Datasets

delete2014-08-25
delete6
delete
OA
AI
C
Carlos Ordońẽz *
C
Carlos Garcia-Alvarado
V
Veerabhadaran Baladandayuthapani
DOI:10.1145/2629617delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Bayesian models are generally computed with Markov Chain Monte Carlo (MCMC) methods. The main disadvantage of MCMC methods is the large number of iterations they need to sample the posterior distributions of model parameters, especially for large datasets. On the other hand, variable selection remains a challenging problem due to its combinatorial search space, where Bayesian models are a promising solution. In this work, we study how to accelerate Bayesian model computation for variable selection in linear regression. We propose a fast Gibbs sampler algorithm, a widely used MCMC method that incorporates several optimizations. We use a Zellner prior for the regression coefficients, an improper prior on variance, and a conjugate prior Gaussian distribution, which enable dataset summarization in one pass, thus exploiting an augmented set of sufficient statistics. Thereafter, the algorithm iterates in main memory. Sufficient statistics are indexed with a sparse binary vector to efficiently compute matrix projections based on selected variables. Discovered variable subsets probabilities, selecting and discarding each variable, are stored on a hash table for fast retrieval in future iterations. We study how to integrate our algorithm into a Database Management System (DBMS), exploiting aggregate User-Defined Functions for parallel data summarization and stored procedures to manipulate matrices with arrays. An experimental evaluation with real datasets evaluates accuracy and time performance, comparing our DBMS-based algorithm with the R package. Our algorithm is shown to produce accurate results, scale linearly on dataset size, and run orders of magnitude faster than the R package.
Keyword:
Sufficient statistics
variable selection
on-line algorithm
MCMC
Gibbs sampler
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

U
university of houston system
学者数:
1.4W
论文数: 1.4W
被引数: 16
U
university of houston
学者数:
9.7K
论文数: 7.9K
被引数: 11
引用论文

引用论文

MicroRNAs and Xenobiotic Toxicity: An Overview
err2020-01-01
err0
errOAAI
errSatheeswaran Balasubramanian; Kanmani Gunasekaran; Saranyadevi Sasidharan; Vignesh Jeyamanickavel Mathan; Ekambaram Perumal
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
DNA starvation/stationary phase protection protein of Helicobacter pylori as a potential immunodominant antigen for infection detection
err2023-02-12
err0
PREAI
errKangle Zhai; Yanan Gong; Lu Sun; Lihua He; Zhijing Xue; Yaming Yang; Mengyang Fang; Jianzhong Zhang
err分享
err收藏
err分享
err收藏
Mixtures of g priors for Bayesian variable selection贝叶斯变量选择的g先验混合
err2012-01-01
err856
PREAI
errLiang, Feng; Paulo, Rui; Molina, German; Clyde, Merlise A.; Berger, Jim O.
err分享
err收藏
学者 查看更多内容