返回
XeFlow: Streamlining Inter-Processor Pipeline Execution for the Discrete CPU-GPU Platform
DOI:10.1109/TC.2020.2968302.png)
摘要
En 中文
Nowadays, GPUs have achieved high throughput computing by running plenty of threads. However, owing to disjoint memory spaces of discrete CPU-GPU systems, exploiting CPU and GPU within a data processing pipeline is a non-trivial issue, which can only be resolved by the coarse-grained workflow of copy-kernel-copy or its variants in essence. There is an underlying bottleneck caused by frequent inter-processor invocations for fine-grained batch sizes. This article presents XeFlow that enables streamlined execution by leveraging hardware mechanisms inside new generation GPUs. XeFlow significantly reduces costly explicit copy and kernel launching within existing fashions. As an alternative, XeFlow introduces persistent operators that continuously process data through shared topics, which establish efficient inter-processor data channels via hardware page faults. Compared with the default copy-kernel-copy method, XeFlow shows up to $2.4\times \!\sim \!3.1\times$2.4x similar to 3.1x performance advantages in both coarse-grained and fine-grained pipeline execution. To demonstrate its potentials, this article also evaluates two GPU-accelerated applications, including data encoding and OLAP query.
Keyword:
CPU-GPU programming
heterogeneous memory system
GPU scheduling
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.8
论文数:
5.4K
被引数:
9.8K
机构
引用论文
Comparison of responses to transmitter candidates at an N-methylaspartate receptor mediated synapse, in slices of rat cerebral cortex
Neuroscience
IF0

