116
H. Iwashita and M. Nakao
Fig. 7 Eight-variable ping-pong latency on PRIMEHPC FX100
• Unless data size exceeds approximately 8 kB, the latency of non-blocking PUT
does not depend on the amount of data. The graph of non-blocking PUT is very
flat, within ±4%, over the range from 8 B to 4 kB.
• MPI eager communication has no effect on non-blocking for latency hiding. The
eager protocol, including the unbuffering process of the receiver, appears not to
be suitable for non-buffering.
• Non-blocking coarray PUT outperforms MPI eager message passing, except for
very fine grain data. The latency of eight-variable non-blocking PUT is −9% to
54% and 18% to 61%, as compared to eight-variable blocking and non-blocking
MPI eager, respectively. At only two plots for 8 B and 16 B, the non-blocking
PUT is 4% and 9% slower than the values for blocking MPI. Otherwise, nonblocking PUT is faster than MPI eager, and the more block size, the larger
difference in the latency.
H. Iwashita and M. Nakao
Fig. 7 Eight-variable ping-pong latency on PRIMEHPC FX100
• Unless data size exceeds approximately 8 kB, the latency of non-blocking PUT
does not depend on the amount of data. The graph of non-blocking PUT is very
flat, within ±4%, over the range from 8 B to 4 kB.
• MPI eager communication has no effect on non-blocking for latency hiding. The
eager protocol, including the unbuffering process of the receiver, appears not to
be suitable for non-buffering.
• Non-blocking coarray PUT outperforms MPI eager message passing, except for
very fine grain data. The latency of eight-variable non-blocking PUT is −9% to
54% and 18% to 61%, as compared to eight-variable blocking and non-blocking
MPI eager, respectively. At only two plots for 8 B and 16 B, the non-blocking
PUT is 4% and 9% slower than the values for blocking MPI. Otherwise, nonblocking PUT is faster than MPI eager, and the more block size, the larger
difference in the latency.
