返回
StaleLearn: Learning Acceleration with Asynchronous Synchronization Between Model Replicas on PIM
DOI:10.1109/TC.2017.2780237.png)
摘要
En 中文
GPU has become popular with a large amount of parallelism found in learning. While the GPU has been effective for many learning tasks, still many GPU learning applications have low execution efficiency due to sparse data. Sparse data induces divergent memory accesses with low locality, thereby consuming a large fraction of execution time transferring data across the memory hierarchy. Although a considerable effort has been devoted to reducing the memory divergence, iterative-convergent learning provides a unique opportunity to achieve full potential in modern GPUs that it allows different threads to continue computation using stale values. In this paper, we propose StaleLearn, a learning acceleration mechanism to reduce the memory divergence overhead of GPU learning by utilizing the stale value tolerance of the iterative-convergent learning. Based on the stale value tolerance, StaleLearn transforms the problem of divergent memory accesses into the synchronization problem by replicating the model and reduces the synchronization overhead by asynchronous synchronization on Processor-in-Memory (PIM). The stale value tolerance enables a clear task decomposition between the GPU and PIM, which can effectively exploit parallelism between PIM and GPU. On average, our approach accelerates representative GPU learning applications by 3.17 times with existing PIM proposals.
Keyword:
Learning acceleration
asynchronous synchronization
GPU memory divergence
stale value tolerance
problem transformation
processor-in-memory
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
IF:
3.8
论文数:
5.4K
被引数:
9.8K
机构
引用论文
Single Nucleotide Polymorphisms inmiR-122Are Associated with the Risk of Hepatocellular Carcinoma in a Southern Chinese Population单核苷酸多态性在miR-122中与南方中国人群肝癌风险相关

