112
H. Iwashita and M. Nakao
Fig. 4 Diagrams for ping-pong codes. (a) Coarray PUT. (b) Coarray GET. (c) MPI send/recv
eager protocol. (d) MPI send/recv Rendezvous protocol
receiver must copy the received data in the local buffer to the target. The larger the
data, the greater the overhead cost. In the rendezvous protocol (d), negotiations,
including remote address notification, are required prior to communication. The
overhead cost is not negligible when the data is small.
The result of the comparison between coarray PUT/GET and MPI message
passing is shown in Fig. 5. As the underlying communication libraries, FJ-RDMA
and MPI-3 are used on FX100 and GASNet. And MPI-3 is used on HA-PACS.
GET (a) and GET (b) use the code without and with the optimization described
in Sect. 3.3.4, respectively. Bandwidth is the communication data size per elapsed
time, and latency is half of the ping-pong elapsed time. The difference between
GET (a) and GET (b) is the compile time optimization level of the coarray translator
described in Sect. 3.3.4.
The following was found regarding coarray PUT/GET communication.
Bandwidth Coarray PUT and GET slightly outperforms MPI rendezvous communication for large data on FJ-RDMA and MPI-3. On FJ-RDMA/FX100 (a),
the bandwidths of PUT and GET (b) are, respectively, +0.1% to +18% and -0.4%
to +9.3% higher than MPI rendezvous in the rendezvous range of 32k through
32M bytes. In addition, on MPI-3 and/or HA-PACS, the bandwidths of PUT
and GET are, respectively, +0.3% to +0.8% and +0.1% to +1.3% higher in the
rendezvous range of 512k through 32M bytes. Based on the runtime log, zerocopy communication was confirmed to have been performed both in PUT and
GET (b) by selecting the DMA scheme described in Sect. 3.3.1.
Précédent

- 119/265

Suivant