248
M. Sato et al.
the XMP local view programming. The library is implemented by using a low-level
communication layer, uTofu API [3], provided by Fujitsu.
For performance evaluation of XMP local view programming, we used CCS
QCD and NTChem-MINI taken from the coarray version of Fiber Miniapp Suite [4,
5].
To run CCS QCD mini-application [6], eight XMP nodes are assigned to one
node, running in a flat XMP mode. The size and conditions are as follows:
• Target data: Class 2 (32 × 32 × 32 × 32) (strong scaling).
• Compiler options: -Kfast, zfill, simd=2.
• Timing region: sum of “Clover + Clover_inv Performance” and “BiCGStab
(CPU: double precision) Performance” of the built-in timing feature.
Figure 2 shows the speedup of the Fugaku, comparing to the performance of
the K computer. The XMP version archives almost same performance of the MPI
version. Note that the reason of the performance degradation of the XMP version on
the K computer is the overhead of allocation for allocatable coarray used as a buffer
for communication. It is improved by removing this overhead by using the uTofu
communication layer.
The NTChem-MINI is a mini-application taken from NTChem [7], a highperformance software package for molecular electronic structure calculation. An
XMP node is assigned to one node, and within a node, BLAS functions are executed
using 48 cores. The size and conditions are set as follows:
• Target data: taxol (strong scaling).
• Compiler options: -Kfast, simd=2.
• Timing region: “RIMP2_Driver” of the built-in timing feature.
Fig. 2 Speedup of CCS QCD on Fugaku and performance comparing to the K computer
Précédent

- 251/265

Suivant