114
H. Iwashita and M. Nakao
However, on GASNet/HA-PACS (c), PUT and GET (b) were only approximately
60% of the bandwidth of MPI rendezvous for a large amount of data. It is
presumed that data copy was caused internally.
Latency On FJ-RDMA (a) and MPI-3 (b) and (d), PUT and GET (b) have larger
(worse) latency than MPI eager communication in the range of ≤16kB on FX100
and ≤256kB on HA-PACS.
Coarray on GASNet (c) behaves differently than other cases on (a), (b), and (d).
Although the latency is larger than that for MPI for all data sizes, the difference
is smaller than in the other cases. At a data size of 8B, the latency of PUT is 2.93
μs and 2.1 times larger than the one of MPI while 5.73μs and 3.7 times larger
for the case of MPI-3 (d).
Effect of GET optimization For all ranges in all cases, GET (a) has a smaller
bandwidth and a larger latency than GET (b). On FJ-RDMA (a), the bandwidth
is 1.41 to 1.85 times improved in the range of 32kB to 32MB by changing
the object code of GET (a) to GET (b). We found GET (a) caused two extra
memory copies. One copy performs the array assignment by the Fortran library,
and the other copy is from the communication buffer to the result variable of the
array function xmpf_coarray_get_generic. The optimization described
in Sect. 3.3.4 eliminated these two data copies.
The large latency of coarray PUT/GET communication is problematic. In the
next subsection, we discuss how this problem should be solved by the compiler and
the programming.
4.2 Non-blocking Communication
For latency hiding, asynchronous and non-blocking features can be expected in
coarray PUT communication. The principle is shown in Fig. 6.
Figure 6a shows the half pattern of the ping-pong PUT communication. Coarray
one-sided communication is basically asynchronous, unless synchronization is
explicitly specified. Therefore, multiple communications without synchronization,
as shown in (b), are closer to actual applications. In addition, coarray one-sided
communication can be optimized using non-blocking communication, as shown in
(c). Blocking and non-blocking communications can be switched with the runtime
environment variable in the current implementation of the Omni compiler. In MPI
message passing, non-blocking communication can be written with MPI_Isend,
MPI_Irecv, and MPI_Wait.
Figure 7 compares blocking/non-blocking coarray PUT and MPI message
passing communications. The two original graphs are the same as those of Fig. 5a.
Four other graphs display the results of the eight-variable ping-pong program, which
repeats the ping phase, sending eight individual variables from one to the other
in order, and, similarly, the pong phase in the opposite direction. Each block size
indicates the size of variables, and latency includes the time for eight variables.
Précédent

- 121/265

Suivant