返回
Functionality-Based Processing-in-Memory Accelerator for Deep Convolutional Neural Networks
DOI:10.1109/ACCESS.2021.3122818.png)
摘要
En 中文
Processing-in-memory (PIM) architectures show the advantage of handling applications that generate complicated memory request patterns; usually, those kinds of memory streams degrade the application's performance in conventional memory hierarchy systems. In particular, deep convolutional neural networks (DCNNs) processing that consists of several functionalities could be highly optimized if PIM cores can extend the processing capability and data accessibility. In this work, we propose a functionality-based PIM accelerator for DCNNs. We design several modules in addition to the conventional PIM system based on a hybrid memory cube (HMC). First, we compose a new buffer module, namely, a shared cache, in which PIM cores are provided DCNN functionalities and pre-trained weights. The PIM cores subsequently enhance computational utilization and data accessibility. Second, an efficient replacement method complements the shared cache to optimize the data miss rate of DCNN processing. Third, we compose dual prefetchers that can deal with DCNN's memory access patterns, thereby reducing the system's overall latency. Fourth, we compose a PIM scheduler for PIM core-level autonomous request control. The PIM scheduler relieves the host processor of significant computational loads, achieving the overall latency of the system and reducing the energy consumption. By the performance evaluation based on the trace-driven HMC simulator, our proposed model improves average latency and bandwidth by 38.9 and 27.9 % with only 18.7 % more energy consumption compared with conventional HMC-based PIM systems. Our system also achieves scalable processing performance because when the DCNN becomes deeper, it processes faster than conventional PIM systems.
Keyword:
Prefetching
Three-dimensional displays
Random access memory
Computer architecture
Feature extraction
Convolutional neural networks
Bandwidth
3D memory
accelerator architectures
artificial intelligence accelerator
computer system
deep neural network
prefetch
processing-in-memory
期刊
IF:
3.6
论文数:
9.8W
被引数:
29.4W
机构
引用论文
Latent dimensions of social anxiety disorder: A re-evaluation of the Social Phobia Inventory (SPIN)社交焦虑障碍的潜在维度: 对社交恐惧症量表 (SPIN) 的重新评估
Gradient-based learning applied to document recognition基于梯度的学习在文档识别中的应用
PROCEEDINGS OF THE IEEE
IF25.9

