Coarrays in the Context of XcalableMP
111
Table 2 Specifications of the computers and evaluation environment
RIKEN RCCS
RIKEN RCCS HOKUSAI
GreatWare
CCS, University of
Tsukuba
The K computer
Fujitsu PRIMEHPC FX100 HA-PACS/TCA
CPU
SPARK64™VIIIfx,
2 GHz, 128 Gflop/s,
8-core, 1 CPU/node
SPARK64™XIfx,
1.975 GHz, 1 CPU/node,
4-SIMD × 32-core
E5-2680 v2 (Ivy Bridge),
10-core, 224 Gflop/s,
2 CPU/node
Memory
16 GB/node,
32 GB/node,
128 GB/node,
Bandwidth 64 GB/s
Bandwidth 480 GB/s
119.4 GB/s
Interconnect Tofu
Tofu2, 12.5 GB/s × 2
InfiniBand FDR, 7 GB/s
Coarray
Omni XcalableMP 1.3.1 Omni XcalableMP 1.3.1
Omni XcalableMP 1.3.1
Fortran
Fujitsu Fortran 2.0.0
Fujitsu Fortran 2.0.0
Intel Fortran 16.0.4
MPI
Fujitu MPI 2.0.0
Fujitu MPI 2.0.0
Intel MPI 5.1.3
Comm. layer Tofu library
Tofu library
GASNet 1.24.2
(IBV-conduit, built with
Intel compilers)
Table 3 Ping-pong codes
4.1 Fundamental Performance
Using the EPCC Fortran Coarray micro-benchmark [6], we evaluated the ping-pong
performance of PUT and GET communications compared with MPI_Send/Recv.
The codes are briefly shown in Table 3.
Corresponding to the codes in Table 3, Fig. 4 shows how data and messages are
exchanged between two images or processes. In coarray PUT (a) and GET (b), interimage synchronization is necessary for each end of the phases to make the passive
image active and to make the active image passive. Whereas in MPI message passing
(c) and (d), such synchronization is not necessary because both processes are always
active. On the other hand, MPI message passing has its own overhead that coarray
PUT/GET does not have. Since the eager protocol (c) does not use RDMA, the
111
Table 2 Specifications of the computers and evaluation environment
RIKEN RCCS
RIKEN RCCS HOKUSAI
GreatWare
CCS, University of
Tsukuba
The K computer
Fujitsu PRIMEHPC FX100 HA-PACS/TCA
CPU
SPARK64™VIIIfx,
2 GHz, 128 Gflop/s,
8-core, 1 CPU/node
SPARK64™XIfx,
1.975 GHz, 1 CPU/node,
4-SIMD × 32-core
E5-2680 v2 (Ivy Bridge),
10-core, 224 Gflop/s,
2 CPU/node
Memory
16 GB/node,
32 GB/node,
128 GB/node,
Bandwidth 64 GB/s
Bandwidth 480 GB/s
119.4 GB/s
Interconnect Tofu
Tofu2, 12.5 GB/s × 2
InfiniBand FDR, 7 GB/s
Coarray
Omni XcalableMP 1.3.1 Omni XcalableMP 1.3.1
Omni XcalableMP 1.3.1
Fortran
Fujitsu Fortran 2.0.0
Fujitsu Fortran 2.0.0
Intel Fortran 16.0.4
MPI
Fujitu MPI 2.0.0
Fujitu MPI 2.0.0
Intel MPI 5.1.3
Comm. layer Tofu library
Tofu library
GASNet 1.24.2
(IBV-conduit, built with
Intel compilers)
Table 3 Ping-pong codes
4.1 Fundamental Performance
Using the EPCC Fortran Coarray micro-benchmark [6], we evaluated the ping-pong
performance of PUT and GET communications compared with MPI_Send/Recv.
The codes are briefly shown in Table 3.
Corresponding to the codes in Table 3, Fig. 4 shows how data and messages are
exchanged between two images or processes. In coarray PUT (a) and GET (b), interimage synchronization is necessary for each end of the phases to make the passive
image active and to make the active image passive. Whereas in MPI message passing
(c) and (d), such synchronization is not necessary because both processes are always
active. On the other hand, MPI message passing has its own overhead that coarray
PUT/GET does not have. Since the eager protocol (c) does not use RDMA, the
