Coarrays in the Context of XcalableMP
121
6 Conclusion
This chapter described the coarray features in the context of XMP and the
characteristic implementation of the coarray translator.
For memory allocation and registration, the RS, RA, and CA methods were
implemented corresponding to the communication library GASNet, FJ-RDMA, and
MPI-3.
For the coarray PUT and GET communications, DMA and four buffering
methods were described. The effect of the non-blocking PUT communication was
analyzed, and the knowledge is used to make the coarray version of the Himeno
benchmark from the original MPI version. The measurement results on 1024 nodes
of the K computer, the coarray version is 27% and 42% faster than the original MPI
version for Himeno sizes L and XL, respectively. The effect of the optimization
of GET communication was also obvious on the ping-pong benchmark on HAPACS/TCA and Fujitsu PRIMEHPC FX100.
As an evaluation of productivity, the coarray program uses fewer than half as
many characters as the MPI message passing program to write the same algorithm
as the Himeno benchmark.
Acknowledgments The present research used the computational resources of HA-PACS provided
by the Interdisciplinary Computational Science Program at the Center for Computational Sciences
at the University of Tsukuba.
The results were obtained in part using the K computer at the RIKEN Advanced Institute for
Computational Science.
References
1. XcalableMP Language Specification, http://xcalablemp.org/specification.html
2. J. Reid, JKR Associates, UK. Coarrays in the next Fortran Standard. ISO/IEC
JTC1/SC22/WG5 N1824, April 21, 2010
3. ISO/IEC TS 18508:2015, Information technology – Additional Parallel Features in Fortran,
Technical Specification, December 1, 2015
4. Omni Compiler Project, http://omni-compiler.org
5. H. Iwashita, M. Nakao, M. Sato, Preliminary implementation of coarray Fortran translator
based on Omni XcalableMP, in PGAS2015, Proceedings of 9th International Conference on
PGAS Programming Models, Washington, DC (2015), pp.70–75
6. EPCC Fortran Coarray micro-benchmark suite, https://www.epcc.ed.ac.uk/research/
computing/performance-characterisation-and-benchmarking/epcc-co-array-fortran-micro
7. Himeno Benchmark, http://accc.riken.jp/en/supercom/himenobmt/
8. J. Mellor-Crummey, L. Adhianto, W.N. Scherer III, G. Jin, PGAS’09, 3rd Conference on
Partitioned Global Address Space Programming Models, Ashburn, VA (2009)
Précédent

- 128/265

Suivant