arrow
返回

CPU cache prefetching: Timing evaluation of hardware implementations

delete1998-05-01
delete26
PRE
AI
T
Tse, J *
S
Smith, AJ
DOI:10.1109/12.677225delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
Prefetching into CPU caches has long been known to be effective in reducing the cache miss ratio, but known implementations of prefetching have been unsuccessful in improving CPU performance. The reasons for this are that prefetches interfere with normal cache operations by making cache address and data ports busy, the memory bus busy, the memory banks busy, and by not necessarily being complete by the time that the prefetched data is actually referenced. In this paper, we present extensive quantitative results of a detailed cycle-by-cycle trace-driven simulation of a uniprocessor memory system in which we vary most of the relevant parameters in order to determine when and if hardware prefetching is useful. We find that, in order for prefetching to actually improve performance, the address array needs to be double ported and the data array needs to either be double ported or fully buffered. It is also very helpful for the bus to be very wide (e.g., 16 bytes) for bus transactions to be split and for main memory to be interleaved. Under the best circumstances, i.e., with a significant investment in extra hardware, prefetching can significantly improve performance. For implementations without adequate hardware, prefetching often decreases performance.
Keyword:
cache memory
prefetching
timing model
cache prefetching
CPU architecture
memory system design
CPU cache memory

期刊

IEEE Transactions on Computers 封面图
IEEE Transactions on Computers
IF:
3.8
论文数:
5.4K
被引数:
9.8K

机构

暂无机构信息
引用论文

引用论文

Strongly asymmetric square waves in a time-delayed system
err2012-11-29
err0
errOAAI
errLionel Weicker; Thomas Erneux; Otti D’Huys; Jan Danckaert; Maxime Jacquot; Yanne Chembo; Laurent Larger
err分享
err收藏
err分享
err收藏
err分享
err收藏
学者 查看更多内容