TY - GEN
T1 - On the multi-GPU computing of a reconstructed discontinuous Galerkin method for compressible flows on 3D hybrid grids
AU - Xia, Yidong
AU - Lou, Jialin
AU - Luo, Lixiang
AU - Luo, Hong
AU - Edwards, Jack
AU - Mueller, Frank
PY - 2014
Y1 - 2014
N2 - A multi-GPU accelerated, third-order, reconstructed discontinuous Galerkin method, namely RDG(P1P2), has been developed based on the OpenACC directives for compressible flows on 3D hybrid grids. The present scheme requires minimum intrusion and algorithm alteration to an existing CPU code, which renders an efficient design approach for upgrading a legacy CFD solver with the GPU-computing capability while maintaining its portability across multiple platforms. The grid partitioning is performed according to the number of GPUs, and loaded equally on each GPU. Communication between the GPUs is achieved via the host-based MPI. A face renumbering and grouping algorithm is used to eliminate memory contention due to vectorized computing over the face loops on each individual GPU. A series of inviscid and viscous flow problems have been presented for the verification and scaling test, demonstrating excellent scalability of the resulting GPU code. The numerical results indicate that this parallel RDG(P1P2) method is a cost-effective, high-order DG method for scalable computing on GPU clusters.
AB - A multi-GPU accelerated, third-order, reconstructed discontinuous Galerkin method, namely RDG(P1P2), has been developed based on the OpenACC directives for compressible flows on 3D hybrid grids. The present scheme requires minimum intrusion and algorithm alteration to an existing CPU code, which renders an efficient design approach for upgrading a legacy CFD solver with the GPU-computing capability while maintaining its portability across multiple platforms. The grid partitioning is performed according to the number of GPUs, and loaded equally on each GPU. Communication between the GPUs is achieved via the host-based MPI. A face renumbering and grouping algorithm is used to eliminate memory contention due to vectorized computing over the face loops on each individual GPU. A series of inviscid and viscous flow problems have been presented for the verification and scaling test, demonstrating excellent scalability of the resulting GPU code. The numerical results indicate that this parallel RDG(P1P2) method is a cost-effective, high-order DG method for scalable computing on GPU clusters.
UR - https://www.scopus.com/pages/publications/85087593744
U2 - 10.2514/6.2014-3081
DO - 10.2514/6.2014-3081
M3 - Conference contribution
AN - SCOPUS:85087593744
SN - 9781624102936
T3 - AIAA AVIATION 2014 -7th AIAA Theoretical Fluid Mechanics Conference
BT - AIAA AVIATION 2014 -7th AIAA Theoretical Fluid Mechanics Conference
PB - American Institute of Aeronautics and Astronautics Inc.
T2 - AIAA AVIATION 2014 -7th AIAA Theoretical Fluid Mechanics Conference 2014
Y2 - 16 June 2014 through 20 June 2014
ER -