arrow
返回

PISA-DMA: Processing-in-Memory Instruction Set Architecture Using DMA

delete2023-01-01
delete2
delete
OA
AI
C
Chang Hyun Kim
Y
Yoonah Paik
S
Seon Wook Kim *
DOI:10.1109/ACCESS.2023.3238812delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Processing-in-memory (PIM) has attracted attention to overcome the memory bandwidth limitation, especially for computing memory-intensive DNN applications. Most PIM approaches use the CPU's memory requests to deliver instructions and operands to the PIM engines, making a core busy and incurring unnecessary data transfer, thus, resulting in significant offloading overhead. DMA can resolve the issue by transferring a high volume of successive data without intervening CPU and polluting the memory hierarchy, thus perfectly fitting the PIM concept. However, the small computing resources of DRAM-based PIM devices allow us to transfer only small amounts of data at one DMA transaction and require a large number of descriptors, thus still incurring significant offloading overhead. This paper introduces PIM Instruction Set Architecture (ISA) using a DMA descriptor called PISA-DMA to express a PIM opcode and operand in a single descriptor. Our ISA makes PIM programming intuitive by thinking of committing one PIM instruction as completing one DMA transaction and representing a sequence of PIM instructions using the DMA descriptor list. Also, PISA-DMA minimizes the offloading overhead while guaranteeing compatibility with commercial platforms. Our PISA-DMA eliminates the opcode offloading overhead and achieves 1.25x, 1.31x, and 1.29x speedup over the baseline PIM at the sequence length of 128 with the BERT, RoBERTa, and GPT-2 models, respectively, in ONNX runtime with real machines. Also, we study how our proposed PISA affects performance in compiler optimization and show that the operator fusion of matrix-matrix multiplication and element-wise addition achieved 1.04x speedup, a similar performance gain using conventional ISAs.
Keyword:
Engines
Computer architecture
Registers
Switches
DRAM chips
Random access memory
Programming
Processing-in-DRAM
direct memory access
instruction set architecture
PIM offloading

期刊

IEEE Access 封面图
IEEE Access
IF:
3.6
论文数:
9.8W
被引数:
29.4W

机构

K
Korea University
学者数:
3.6W
论文数: 3.8W
被引数: 4.4W
引用论文

引用论文

Interhemispheric co-alteration of brain homotopic regions半球间大脑同源区域的协同改变
err2021-06-25
err0
errOAAI
errFranco Cauda; Andrea Nani; Donato Liloia; Gabriele Gelmini; Lorenzo Mancuso; Jordi Manuello; Melissa Panero; Sergio Duca; Yu-Feng Zang; Tommaso Costa
err分享
err收藏
Supported metal complex catalysts
err1975-09-01
err0
PREAI
errPeter R. Rony; James F. Roth
err分享
err收藏
err分享
err收藏
err分享
err收藏
Periconceptional undernutrition in sheep leads to decreased locomotor activity in a natural environment
err2013-05-31
err0
PREAI
errE. L. Donovan; C. E. Hernandez; L. R. Matthews; M. H. Oliver; A. L. Jaquiery; F. H. Bloomfield; J. E. Harding
err分享
err收藏
学者 查看更多内容