arrow
返回

Efficient spatial data partitioning for distributed kNN joins

delete2022-06-02
delete2
delete
OA
AI
A
Ayman Zeidan *
H
Huy T. Vo
DOI:10.1186/s40537-022-00587-2delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Parallel processing of large spatial datasets over distributed systems has become a core part of modern data analytic systems like Apache Hadoop and Apache Spark. The general-purpose design of these systems does not natively account for the data's spatial attributes and results in poor scalability, accuracy, or prolonged runtimes. Spatial extensions remedy the problem and introduce spatial data recognition and operations. At the core of a spatial extension, a locality-preserving spatial partitioner determines how to spatially group the dataset's objects into smaller chunks using the distributed system's available resources. Existing spatial extensions rely on data sampling and often mismanage non-spatial data by either overlooking their memory requirements or excluding them entirely. This work discusses the various challenges that face spatial data partitioning and proposes a novel spatial partitioner for effectively processing spatial queries over large spatial datasets. For evaluation, the proposed partitioner is integrated with the well-known k-Nearest Neighbor (kNN) spatial join query. Several experiments evaluate the proposal using real-world datasets. Our approach differs from existing proposals by (1) accounting for the dataset's unique spatial traits without sampling, (2) considering the computational overhead required to handle non-spatial data, (3) minimizing partition shuffles, (4) computing the optimal utilization of the available resources, and (5) achieving accurate results. This contributes to the problem of spatial data partitioning through (1) providing a comprehensive discussion of the problems facing spatial data partitioning and processing, (2) the development of a novel spatial partitioning technique for in-memory distributed processing, (3) an effective, built-in, load-balancing methodology that reduces spatial query skews, and (4) a Spark-based implementation of the proposed work with an accurate kNN spatial join query. Experimental tests show up to 1.48 times improvement in runtime as well as the accuracy of results.
Keyword:
Big data
Spatial data
Spatial query
Indexing
Partitioning
Technique
Parallel processing
Load balancing
Distributed computing
Spark
NoSQL
kNN query
All kNN query

期刊

Journal of Big Data 封面图
Journal of Big Data
IF:
6.4
论文数:
1.5K
被引数:
1.1W

机构

C
city university of new york (cuny) system
学者数:
1.6W
论文数: 1.5W
被引数: 26
引用论文

引用论文

Structural Features and Physical Properties of In2Bi3Se7I, InBi2Se4I, and BiSeI
err2011-10-27
err0
PREAI
errTobias Rosenthal; Markus Döblinger; Peter Wagatha; Christian Gold; Ernst‐Wilhelm Scheidt; Wolfgang Scherer; Oliver Oeckler
err分享
err收藏
Transcriptome Sequencing and Expression Analysis of Terpenoid Biosynthesis Genes in Litsea cubeba
err2013-10-09
err0
errOAAI
errXiao-Jiao Han; Yang-Dong Wang; Yi-Cun Chen; Li-Yuan Lin; Qing-Ke Wu
err分享
err收藏
Multimodal control for human-robot cooperation
err2013-11-01
err0
errOAAI
errAndrea Cherubini; Robin Passama; Arnaud Meline; Andre Crosnier; Philippe Fraisse
err分享
err收藏
err分享
err收藏
Comparison of the springtime vertical export of biogenic matter in three northern Norwegian fjords
err2000-01-01
err0
errOAAI
errM Reigstad; P Wassmann; T Ratkova; E Arashkevich; A Pasternak; S Øygarden
err分享
err收藏
Excitonic Dark States in Single Atomic Layer of Transition Metal Dichalcogenide
err2014-01-01
err0
PREAI
errZiliang Ye; Ting Cao; Kevin O’Brien; Hanyu Zhu; Xiaobo Yin; Yuan Wang; Steven G. Louie; Xiang Zhang
err分享
err收藏
err分享
err收藏
Life History of the Mottled Scorpionfish, Pontinus clemensi, in the Galapagos Marine Reserve
err2018-10-01
err0
PREAI
errJ. R. Marin Jarrin; S. Andrade-Vera; C. Reyes-Ojedis; P. Salinas-de-León
err分享
err收藏
学者 查看更多内容