arrow
Return

High-Performance, Graphics Processing Unit-Accelerated Fock Build Algorithm

delete2020-11-18
delete35
PRE
AI
G
Giuseppe M. J. Barca *
J
Jorge L. Galvez-Vallejo
D
David Poole
A
Alistair P. Rendell
M
Mark S. Gordon
DOI:10.1021/acs.jctc.0c00768delete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
We present a high-performance, GPU (graphics processing unit)-accelerated algorithm for building the Fock matrix. The algorithm is designed for efficient calculations on large molecular systems and uses a novel dynamic load balancing scheme that maximizes the GPU throughput and avoids thread divergence that could occur due to integral screening. Additionally, the code adopts a novel ERI digestion algorithm that exploits all forms of perfmutational symmetry, combines efficiently the evaluation of both Coulomb and exchange terms together, and eliminates explicit thread synchronization requirements. Performance results obtained using a number of large molecules reveal remarkable speedups up to 24.4x with respect to the QUICK GPU code and up to 237x with respect to the GAMESS CPU parallel code.
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Journal of Chemical Theory and Computation cover
Journal of Chemical Theory and Computation
IF:
5.5
Papers:
1.1W
Citations:
5.4W

Organization

I
Iowa State University
Scholars:
2.1W
Papers: 1.8W
Citations: 2.5W
A
Ames National Laboratory
Scholars:
1.3K
Papers: 904
Citations: 3.0K
U
united states department of energy (doe)
Scholars:
11.3W
Papers: 9.6W
Citations: 246
researcher View more organizations