arrow
返回

Energy-Efficient Floating-Point Unit Design

delete2011-07-01
delete100
PRE
AI
S
Sameh Galal *
M
Mark Horowitz
DOI:10.1109/TC.2010.121delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Energy-efficient computation is critical if we are going to continue to scale performance in power-limited systems. For floating-point applications that have large amounts of data parallelism, one should optimize the throughput/mm(2) given a power density constraint. We present a method for creating a trade-off curve that can be used to estimate the maximum floating-point performance given a set of area and power constraints. Looking at FP multiply-add units and ignoring register and memory overheads, we find that in a 90 nm CMOS technology at 1 W/mm(2), one can achieve a performance of 27 GFlops/mm(2) single precision, and 7.5 GFlops/mm(2) double precision. Adding register file overheads reduces the throughput by less than 50 percent if the compute intensity is high. Since the energy of the basic gates is no longer scaling rapidly, to maintain constant power density with scaling requires moving the overall FP architecture to a lower energy/performance point. A 1 W/mm(2) design at 90 nm is a high-energy design, so scaling it to a lower energy design in 45 nm still yields a 7 x performance gain, while a more balanced 0.1 W/mm(2) design only speeds up by 3.5 x when scaled to 45 nm. Performance scaling below 45 nm rapidly decreases, with a projected improvement of only similar to 3 x for both power densities when scaling to a 22 nm technology.
Keyword:
Arithmetic and logic structures
high-speed arithmetic
floating point
fused multiply-add
throughput/mm(2) optimization

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.4K
被引数:
9.8K

机构

S
Stanford University
学者数:
9.6W
论文数: 8.2W
被引数: 17.0W
引用论文

引用论文

Magnetic, structural, and Raman characterization ofRBa2Cu2NbO8(R=Pr, La, or Nd)
err1992-11-01
err0
PREAI
errM. Bennahmias; J. C. O’Brien; H. B. Radousky; T. J. Goodwin; P. Klavins; J. M. Link; C. A. Smith; R. N. Shelton
err分享
err收藏
GPU computingGPU计算
err2008-05-01
err1.4K
PREAI
errOwens, John D.; Houston, Mike; Luebke, David; Green, Simon; Stone, John E.; Phillips, James C.
err分享
err收藏
Refinement and reduction in production of genetically modified mice
err2003-07-01
err0
errOAAI
errV. Robinson; D. B. Morton; D. Anderson; J. F. A. Carver; R. J. Francis; R. Hubrecht; E. Jenkins; K. E. Mathers; R. Raymond; I. Rosewell; J. Wallace; D. J. Wells
err分享
err收藏
Broadly tunable (993–1110  nm) Yb:YLF laser
err2022-04-21
err0
errOAAI
errUmit Demirbas; Jelto Thesinga; Martin Kellert; Simon Reuter; Mikhail Pergament; Franz X. Kärtner
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
High field effects in high resistivity silicon carbide in lateral configurations
err1995-12-04
err0
PREAI
errT. S. Sudarshan; G. Gradinaru; G. Korony; W. Mitchel; R. H. Hopkins
err分享
err收藏
没有更多内容