arrow
返回

An NVM SSD-Based High Performance Query Processing Framework for Search Engines

delete2022-01-01
delete1
PRE
AI
X
Xinyu Liu
Y
Yu Pan
Y
Yusen Li
王刚 封面图
王刚 (Gang Wang)
DOI:10.1109/TKDE.2022.3160557delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Commercial search engines generally maintain hundreds of thousands of machines equipped with large sized DRAM in order to process huge volume of user queries with fast responsiveness, which incurs high hardware cost since DRAM is very expensive. Recently, NVM Optane SSD has been considered as a promising underlying storage device due to its price advantage over DRAM and speed advantage over traditional slow block devices. However, to achieve a comparable efficiency performance with in-memory index, applying NVM to both latency and I/O bandwidth critical applications such as search engines still face non-trivial challenges, because NVM has much lower I/O speed and bandwidth compared to DRAM. In this paper, we propose an NVM SSD-optimized query processing framework, aiming to address both the latency and bandwidth issues of using NVM in search engines. First, we propose a pipelined query processing methodology which significantly reduces the I/O waiting time by fine-grained overlapping of the computation and I/O operations. Second, we propose a cache-aware query reordering algorithm which enables queries sharing more data to be processed adjacently so that the I/O traffic is minimized. Third, we propose a data prefetching mechanism which reduces the extra thread waiting time due to data sharing and improves bandwidth utilization. Moreover, we propose intra-query parallel mechanisms for long-tail queries, including query subtask scheduling, heap concurrent access strategy, query parallelism prediction and adaptive pipelining. Extensive experimental studies show that our framework significantly outperforms the state-of-the-art baselines, which obtains comparable processing latency and throughput with DRAM while using much less space in both inter-query and intra-query parallel scenarios.
Keyword:
Random access memory
Bandwidth
Search engines
Indexes
Nonvolatile memory
Query processing
Throughput
NVM Optane SSD
inter-query and intra-query parallel processing
pipeline
query reordering
prefetching

期刊

IEEE Transactions on Knowledge and Data Engineering 封面图
IEEE Transactions on Knowledge and Data Engineering
IF:
10.4
论文数:
6.8K
被引数:
3.2W

机构

N
nankai university
学者数:
4.8W
论文数: 3.3W
被引数: 74
引用论文

引用论文

err分享
err收藏
Chemical cross‐linking of the chloroplast localized small heat‐shock protein, Hsp21, and the model substrate citrate synthase
err2009-01-02
err0
errOAAI
errEmma Åhrman; Wietske Lambert; J. Andrew Aquilina; Carol V. Robinson; Cecilia Sundby Emanuelsson
err分享
err收藏
Fast top-k preserving query processing using two-tier indexes
err2016-09-01
err16
PREAI
errDaoud, Caio Moura; de Moura, Edleno Silva; Carvalho, Andre; da Silva, Altigran Soares; Fernandes, David; Rossi, Cristian
err分享
err收藏
The Poincaré-sphere approach to polarization: Formalism and new labs with Poincaré beams
err2016-11-01
err0
PREAI
errJoshua A. Jones; Anthony J. D’Addario; Brett L. Rojec; G. Milione; Enrique J. Galvez
err分享
err收藏
err分享
err收藏
学者 查看更多内容