118
H. Iwashita and M. Nakao
number of communications increases from four to eight per node, but all
communications become independent and can be non-blocking.
Coarray PUT/non-blocking The communication pattern is the same as that of
MPI/non-blocking. The data was declared as a coarray, and each communication
was written with a coarray PUT, i.e., an assignment statement with coindexed
variable as the left-hand side. Since the right-hand side of the statement is a reference to the same variable as the left-hand side coarray, the PUT communication
was converted to zero-copy DMA-RDMA communication.
4.3.2 Measurement Result
Figure 9 shows the measurement results for Himeno sizes M, L, and XL, executed
on 1×1, 2×2, 4×4, · · · , 32×32 nodes on the K computer. The following results
were obtained:
• PUT non-blocking was the fastest at 76% of the measurement points of the
graph. On 1024 nodes, PUT non-blocking is 1.2%, 27%, and 42% faster than
MPI original for sizes M, L, and XL, respectively.
• As a result of analyzing the contents of elapsed time, it was confirmed that the
difference in the performance is caused by the difference in communication time.
As shown in (b) and (c), the communication times of PUT non-blocking are 56%
and 51% of those of MPI blocking on 256 nodes on L and XL Himeno sizes,
respectively.
• MPI non-blocking is not always faster than the MPI original. The effect of nonblocking seems to be limited in MPI.
4.3.3 Productivity
Table 4 compares the scale of the source codes. The following features can be found.
• PUT blocking requires fewer characters for programming, especially in subroutines initcomm and initmax. The MPI programmer must describe the
Cartesian coordinates to represent neighboring nodes in initcomm, and must
declare MPI vector types to describe the communication pattern in initmax.
In contrast, the coarray programmer easily represents neighboring images with
coindex notation, e.g., [i,j-1,k], and communication patterns with subarray
notations, e.g., p(1:imax, 1:jmax, 1).
• The Fortran statement of the MPI program tends to be longer than that of
coarray. Since, comparing PUT non-blocking to MPI non-blocking, the number
of characters is one third while the number of statements is almost the same. This
means that the coarray program is more compact than the MPI program for each
statement. MPI library functions often require long sequences of arguments.
Précédent

- 125/265

Suivant