XcalableMP 2.0 and Future Directions
249
Fig. 3 Speedup of NTChem-MINI on Fugaku and performance comparing to the K computer
As shown in Fig. 3, the XMP versions archive almost the same performance of
the original MPI versions.
3 Global Task Parallel Programming
Recently, large-scale clusters of manycore processors such as Intel Xeon Phi have
been deployed in many sites from the latest Top500 Lists. In order to program
manycore processors, OpenMP is widely used as a shared-memory programming
model. Most OpenMP programs are written using work sharing constructs for
loops, which involves a global synchronization. However, especially in modern
manycore processors, the global synchronization cost for work sharing becomes
bigger, and the load imbalance among cores lead to the performance degradation
as the number of cores on the processor increases. Task parallel programming
using task dependency in OpenMP 4.0 is a promising candidate to facilitate the
parallelization for such manycore processors because it enables users to avoid global
synchronization by fine-grained task-to-task synchronization through user-specified
data dependencies.
We are interested in extending the task parallel programming model to the
PGAS model of XcalableMP for distributed memory systems. As well as removing
expensive global synchronization, it is expected to enable the overlapping of
communication and computation. For XMP 2.0, we propose the global task parallel
programming.
In OpenMP, the task dependency in a node depends on the order of reading and
writing to data based on the sequential execution. Therefore, the OpenMP multi-
Précédent

- 252/265

Suivant