XcalableMP 2.0 and Future Directions
247
The performance of XcalableMP on the Fugaku is enhanced by the manycore
processor and a new Tofu-D interconnect.
2.1 Performance of XcalableMP Global View Programming
We executed the IMPACT-3D, described in Chap. 6, for the evaluation of XcalableMP global view programming in the Fugaku, using up to 512 nodes. The
scalability on Fugaku is shown in Fig. 1, comparing to the MPI version. The program
is parallelized by hybrid XMP-OpenMP parallel programming: An XMP node is
assigned to a node, and 48 OpenMP threads are running within a node. The problem
size is 512 × 512 × 512 with three-dimensional block distribution. The compile
option is “-Kfast”.
As shown in the figure, we found a good scalability in Fugaku, and the
performance is better than that by MPI thanks to the optimized XMP runtime for
communications in the stencil computation [2].
2.2 Performance of XcalableMP Local View Programming
Fugaku has a customized interconnection, called Tofu-D, which provides hardwaresupported RDMA (Remote Direct Memory Access) operations. We implemented
the XMP runtime library to make use of Tofu-D for one-sided communication for
Fig. 1 Speedup of Impact3D on Fugaku and performance comparing to K computer
247
The performance of XcalableMP on the Fugaku is enhanced by the manycore
processor and a new Tofu-D interconnect.
2.1 Performance of XcalableMP Global View Programming
We executed the IMPACT-3D, described in Chap. 6, for the evaluation of XcalableMP global view programming in the Fugaku, using up to 512 nodes. The
scalability on Fugaku is shown in Fig. 1, comparing to the MPI version. The program
is parallelized by hybrid XMP-OpenMP parallel programming: An XMP node is
assigned to a node, and 48 OpenMP threads are running within a node. The problem
size is 512 × 512 × 512 with three-dimensional block distribution. The compile
option is “-Kfast”.
As shown in the figure, we found a good scalability in Fugaku, and the
performance is better than that by MPI thanks to the optimized XMP runtime for
communications in the stencil computation [2].
2.2 Performance of XcalableMP Local View Programming
Fugaku has a customized interconnection, called Tofu-D, which provides hardwaresupported RDMA (Remote Direct Memory Access) operations. We implemented
the XMP runtime library to make use of Tofu-D for one-sided communication for
Fig. 1 Speedup of Impact3D on Fugaku and performance comparing to K computer
