92
M. Nakao and H. Murai
step. Therefore, the chunk size is about 4096 Bytes (= 1024/2 × 64 bits/8). Note
that the destination node cannot know how many elements are sent by the source
node. Thus, the MPI implementation gets the number of elements using the function
MPI_Get_count(). We implement the algorithm using a coarray and the post/wait
directives for the recursive exchange algorithm, and the number of elements is added
to the first element of the coarray.
Figure 17 shows a part of the XMP implementation. In line 2, the coarrays
recv[][][] and send[][] are declared. In line 6, the data chunk size is set at the first
element of the coarray, and it is put in line 7. In line 8, the node sends notification of
the completion of the coarray operation of line 7 to the node p[ipartner]. In line 10,
the node receives the notification from the node p[jpartner], which ensures that the
node p[jpartner] receives the data. In line 11, the node gets the number of elements
in the received data. In line 12, the node updates own table by using the received
data.
5.5.3 Evaluation
Figure 18 shows the performance results and performance ratios. The Giga-updates
per second (GUPS) on the vertical axis is the measurement value, which is the
number of update tables per second divided by 10 9 . XMP’s best performance results
are 259.73 GUPS for 16,384 compute nodes on the K computer, and 6.23 GUPS for
128 compute nodes on the COMA system. The values of the performance ratio are
between 1.01 and 1.11 on the K computer, and between 0.57 and 1.03 on the COMA
system. On the K computer, the performance results for the XMP implementation
are always slightly better than those for the MPI implementation. However, on the
COMA system, the performance results for the XMP implementation are worse than
those for the MPI implementation using multiple CPUs.
1 #pragma xmp nodes p[*]
2 unsigned long long recv[ITER][LOGPROCS][CHUNK]:[*], send[2][CHUNKBIG]:[*];
3 ...
4 for(j=0;j
5
...
6
send[i][0] = nsend;
7
recv[iter_mod][j][0:nsend+1]:[ipartner] = send[i][0:nsend+1];
8 #pragma xmp post(p[ipartner], tag)
9
...
10 #pragma xmp wait(p[jpartner], tag)
11
nrecv = recv[iter_mod][j−1][0];
12
update_table(&recv[iter_mod][j−1][1], ..., nrecv, ...);
13
...
14 }
Fig. 17 Part of the RandomAccess code [5]
M. Nakao and H. Murai
step. Therefore, the chunk size is about 4096 Bytes (= 1024/2 × 64 bits/8). Note
that the destination node cannot know how many elements are sent by the source
node. Thus, the MPI implementation gets the number of elements using the function
MPI_Get_count(). We implement the algorithm using a coarray and the post/wait
directives for the recursive exchange algorithm, and the number of elements is added
to the first element of the coarray.
Figure 17 shows a part of the XMP implementation. In line 2, the coarrays
recv[][][] and send[][] are declared. In line 6, the data chunk size is set at the first
element of the coarray, and it is put in line 7. In line 8, the node sends notification of
the completion of the coarray operation of line 7 to the node p[ipartner]. In line 10,
the node receives the notification from the node p[jpartner], which ensures that the
node p[jpartner] receives the data. In line 11, the node gets the number of elements
in the received data. In line 12, the node updates own table by using the received
data.
5.5.3 Evaluation
Figure 18 shows the performance results and performance ratios. The Giga-updates
per second (GUPS) on the vertical axis is the measurement value, which is the
number of update tables per second divided by 10 9 . XMP’s best performance results
are 259.73 GUPS for 16,384 compute nodes on the K computer, and 6.23 GUPS for
128 compute nodes on the COMA system. The values of the performance ratio are
between 1.01 and 1.11 on the K computer, and between 0.57 and 1.03 on the COMA
system. On the K computer, the performance results for the XMP implementation
are always slightly better than those for the MPI implementation. However, on the
COMA system, the performance results for the XMP implementation are worse than
those for the MPI implementation using multiple CPUs.
1 #pragma xmp nodes p[*]
2 unsigned long long recv[ITER][LOGPROCS][CHUNK]:[*], send[2][CHUNKBIG]:[*];
3 ...
4 for(j=0;j
...
6
send[i][0] = nsend;
7
recv[iter_mod][j][0:nsend+1]:[ipartner] = send[i][0:nsend+1];
8 #pragma xmp post(p[ipartner], tag)
9
...
10 #pragma xmp wait(p[jpartner], tag)
11
nrecv = recv[iter_mod][j−1][0];
12
update_table(&recv[iter_mod][j−1][1], ..., nrecv, ...);
13
...
14 }
Fig. 17 Part of the RandomAccess code [5]
