Coarrays in the Context of XcalableMP
Hidetoshi Iwashita and Masahiro Nakao
Abstract Coarray features have been implemented on the Omni XcalableMP
compiler with a source-to-source translator and layered runtime libraries. Three
memory allocation methods for coarrays were implemented for the GASNet and
MPI-3 communication libraries and the native interface of Fujitsu. For the coarray PUT/GET communication, algorithms using DMA (zero-copy) and buffering
were introduced. Important techniques for achieving high performance were the
non-blocking PUT communication implemented in the runtime library and the
optimization for the GET communication in the translator. Using the ping-pong
benchmark and the modified version, the fundamental performance was evaluated
and analyzed. The MPI version of the Himeno benchmark was ported to the coarray
version and modified for fully using the non-blocking PUT. As a result of the
evaluation, the non-blocking coarray version clearly outperformed the original and
non-blocking MPI versions.
1 Introduction
XcalableMP (XMP) [1] has complementary global-view and local-view programming models. The former is a directive-based language extension to the base
languages Fortran and C, and the latter adopts the coarray features defined in Fortran
2008 [2] and a part of the coarray features defined in Fortran 2018 [3]. The purpose
of the coarray features as the local-view part of XMP is (1) writing applications
that are not suitable for global-view programming and (2) writing important parts of
programs that are critical to performance with an easier programming model than
MPI message passing. Therefore, the coarray features in XMP must be naturally
H. Iwashita ()
Fujitsu Limited, Numazu-shi, Shizuoka, Japan
e-mail: iwashita.hideto@fujitsu.com
M. Nakao
RIKEN Center for Computational Science, Kobe, Hyogo, Japan
e-mail: masahiro.nakao@riken.jp
© The Author(s) 2021
M. Sato (ed.), XcalableMP PGAS Programming Language,
https://doi.org/10.1007/978-981-15-7683-6_3
97
Hidetoshi Iwashita and Masahiro Nakao
Abstract Coarray features have been implemented on the Omni XcalableMP
compiler with a source-to-source translator and layered runtime libraries. Three
memory allocation methods for coarrays were implemented for the GASNet and
MPI-3 communication libraries and the native interface of Fujitsu. For the coarray PUT/GET communication, algorithms using DMA (zero-copy) and buffering
were introduced. Important techniques for achieving high performance were the
non-blocking PUT communication implemented in the runtime library and the
optimization for the GET communication in the translator. Using the ping-pong
benchmark and the modified version, the fundamental performance was evaluated
and analyzed. The MPI version of the Himeno benchmark was ported to the coarray
version and modified for fully using the non-blocking PUT. As a result of the
evaluation, the non-blocking coarray version clearly outperformed the original and
non-blocking MPI versions.
1 Introduction
XcalableMP (XMP) [1] has complementary global-view and local-view programming models. The former is a directive-based language extension to the base
languages Fortran and C, and the latter adopts the coarray features defined in Fortran
2008 [2] and a part of the coarray features defined in Fortran 2018 [3]. The purpose
of the coarray features as the local-view part of XMP is (1) writing applications
that are not suitable for global-view programming and (2) writing important parts of
programs that are critical to performance with an easier programming model than
MPI message passing. Therefore, the coarray features in XMP must be naturally
H. Iwashita ()
Fujitsu Limited, Numazu-shi, Shizuoka, Japan
e-mail: iwashita.hideto@fujitsu.com
M. Nakao
RIKEN Center for Computational Science, Kobe, Hyogo, Japan
e-mail: masahiro.nakao@riken.jp
© The Author(s) 2021
M. Sato (ed.), XcalableMP PGAS Programming Language,
https://doi.org/10.1007/978-981-15-7683-6_3
97
