返回
Accelerating the similarity self-join using the GPU
DOI:10.1016/j.jpdc.2019.06.005.png)
摘要
En 中文
The self-join finds all objects in a dataset within a threshold of each other defined by a similarity metric. As such, the self-join is a fundamental building block for the field of databases and data mining. In low dimensionality, there are several challenges associated with efficiently computing the self-join on the graphics processing unit (GPU). Low dimensional data results in higher data densities, causing a significant number of distance calculations and a large result set, and as dimensionality increases, index searches become increasingly exhaustive. We propose several techniques to optimize the self-join using the GPU that include a CPU-efficient index that employs a bounded search, a batching scheme to accommodate large result sets, and duplicate search removal with low overhead. Furthermore, we propose a performance model that reveals bottlenecks related to the result set size and enables us to choose a batch size that mitigates two sources of performance degradation. Our approach outperforms the state-of-the-art on most scenarios. (C) 2019 Elsevier Inc. All rights reserved.
Keyword:
GPGPU
In-memory database
Query optimization
Self-join
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
4
论文数:
3.8K
被引数:
4.8K
机构
引用论文
A Performance Modeling and Optimization Analysis Tool for Sparse Matrix-Vector Multiplication on GPUs一种基于gpu的稀疏矩阵向量乘性能建模与优化分析工具
THE ELEVENTH AND TWELFTH DATA RELEASES OF THE SLOAN DIGITAL SKY SURVEY: FINAL DATA FROM SDSS-III斯隆数字天空调查的第十一和第十二次数据发布: 来自sdss-iii的最终数据


