arrow
返回

Multi-core and many-core shared-memory parallel raycasting volume rendering optimization and tuning

delete2012-04-03
delete15
delete
OA
AI
E
E. Wes Bethel *
M
Mark Howison
DOI:10.1177/1094342012440466delete
delete原文链接
delete分享
delete收藏
查看原文
摘要

摘要

En 中文
Given the computing industry trend of increasing processing capacity by adding more cores to a chip, the focus of this work is tuning the performance of a staple visualization algorithm, raycasting volume rendering, for shared-memory parallelism on multi-core CPUs and many-core GPUs. Our approach is to vary tunable algorithmic settings, along with known algorithmic optimizations and two different memory layouts, and measure performance in terms of absolute runtime and L2 memory cache misses. Our results indicate there is a wide variation in runtime performance on all platforms, as much as 254% for the tunable parameters we test on multi-core CPUs and 265% on many-core GPUs, and the optimal configurations vary across platforms, often in a non-obvious way. For example, our results indicate the optimal configurations on the GPU occur at a crossover point between those that maintain good cache utilization and those that saturate computational throughput. This result is likely to be extremely difficult to predict with an empirical performance model for this particular algorithm because it has an unstructured memory access pattern that varies locally for individual rays and globally for the selected viewpoint. Our results also show that optimal parameters on modern architectures are markedly different from those in previous studies run on older architectures. In addition, given the dramatic performance variation across platforms for both optimal algorithm settings and performance results, there is a clear benefit for production visualization and analysis codes to adopt a strategy for performance optimization through auto-tuning. These benefits will likely become more pronounced in the future as the number of cores per chip and the cost of moving data through the memory hierarchy both increase.
Keyword:
parallel volume rendering
performance optimization
auto-tuning
multi-core CPU
many-core GPU
AI总结

AI总结

对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。

期刊

International Journal of High Performance Computing Applications 封面图
International Journal of High Performance Computing Applications
IF:
2.5
论文数:
1.1K
被引数:
1.3K

机构

L
Lawrence Berkeley National Laboratory
学者数:
1.5W
论文数: 1.1W
被引数: 6.1W
U
united states department of energy (doe)
学者数:
11.3W
论文数: 9.6W
被引数: 246
引用论文

引用论文

Phosphorus Losses through Agricultural Tile Drainage in Nova Scotia, Canada
err2007-03-01
err0
PREAI
errRobert D. Kinley; Robert J. Gordon; Glenn W. Stratton; Gary T. Patterson; Jeff Hoyle
err分享
err收藏
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
Investigation of hydriding properties of LaNi4.8Sn0.2, LaNi4.27Sn0.24 and La0.9Gd0.1Ni5 after thermal cycling and aging
err1992-08-01
err0
PREAI
errSteven W. Lambert; Dhanesh Chandra; William N. Cathey; Franklin E. Lynch; Robert C. Bowman
err分享
err收藏
学者 查看更多内容