Coarrays in the Context of XcalableMP
107
3.3.2 Buffering Communication Methods
For the buffer scheme, one of the four algorithms will be chosen depending on
three parameters: the size of the local buffer B and the local and remote contiguous
lengths N L and N R , respectively. Here, B should be large enough to ignore
communication latency overhead and we use approximately 400 kilo-bytes by
default. Unlike the case of MPI message passing, coarray PUT/GET communication
requires only one local buffer for any number of other images. Both N L and N R can
be evaluated at runtime. The Fortran syntax guarantees that N L is a multiple of N R
or N R is a multiple of N L . An algorithm to obtain the contiguous length is shown in
a previous paper [5].
Table 1 summarizes our algorithm for PUT/GET communication for five cases.
The unit size is the chunk length of the PUT/GET communication. Case 0 shows the
algorithm using RDMA-DMA PUT/GET communication, and Cases 1 through 4
show the algorithms using RDMA and local-buffering. Due to its strict condition,
the DMA scheme is rarely used. In addition, this scheme is not always faster than
the buffering scheme for Cases 2 and 3 because of the difference in the unit sizes.
The advantage of Cases 2 and 3 is that the unit size is extended to a multiple of N L
by gathering a number of short contiguous data in the buffer, or by scattering from
the buffer into a number of short contiguous data.
3.3.3 Non-blocking PUT Communication
For higher performance, the PUT communication should be non-blocking, and the
completion wait should be delayed until the end of the segment. Writing and reading
the same remote data from the same image in the same segment appears to be a
very rare case, as described in Sect. 2.4. However, this is difficult to detect with
Table 1 Summary of the PUT/GET algorithm related to N L , N R , and B
Scheme
Case
Condition
Unit size
DMA
Local data is registered
min(N L , N R )
Buffering
1
N R ≤ B, N R ≤ N L
N R
2
N L < N R ≤ B
N R
3
N L < B < N R
Multiple of N L (≤ B)
4
B < N R , B ≤ N L
B (or less than B at last)
Scheme
Case PUT action for each unit
GET action for each unit
DMA
Put once
Get once
Buffering 1
Buffer once, and put once
Get once, and unbuffer once
2
Buffer for each N L , and put once Get once, and unbuffer for each N L
3
Buffer for each N L , and put once Get once, and unbuffer for each N L
4
Buffer once, and put once
Get once, and unbuffer once
107
3.3.2 Buffering Communication Methods
For the buffer scheme, one of the four algorithms will be chosen depending on
three parameters: the size of the local buffer B and the local and remote contiguous
lengths N L and N R , respectively. Here, B should be large enough to ignore
communication latency overhead and we use approximately 400 kilo-bytes by
default. Unlike the case of MPI message passing, coarray PUT/GET communication
requires only one local buffer for any number of other images. Both N L and N R can
be evaluated at runtime. The Fortran syntax guarantees that N L is a multiple of N R
or N R is a multiple of N L . An algorithm to obtain the contiguous length is shown in
a previous paper [5].
Table 1 summarizes our algorithm for PUT/GET communication for five cases.
The unit size is the chunk length of the PUT/GET communication. Case 0 shows the
algorithm using RDMA-DMA PUT/GET communication, and Cases 1 through 4
show the algorithms using RDMA and local-buffering. Due to its strict condition,
the DMA scheme is rarely used. In addition, this scheme is not always faster than
the buffering scheme for Cases 2 and 3 because of the difference in the unit sizes.
The advantage of Cases 2 and 3 is that the unit size is extended to a multiple of N L
by gathering a number of short contiguous data in the buffer, or by scattering from
the buffer into a number of short contiguous data.
3.3.3 Non-blocking PUT Communication
For higher performance, the PUT communication should be non-blocking, and the
completion wait should be delayed until the end of the segment. Writing and reading
the same remote data from the same image in the same segment appears to be a
very rare case, as described in Sect. 2.4. However, this is difficult to detect with
Table 1 Summary of the PUT/GET algorithm related to N L , N R , and B
Scheme
Case
Condition
Unit size
DMA
Local data is registered
min(N L , N R )
Buffering
1
N R ≤ B, N R ≤ N L
N R
2
N L < N R ≤ B
N R
3
N L < B < N R
Multiple of N L (≤ B)
4
B < N R , B ≤ N L
B (or less than B at last)
Scheme
Case PUT action for each unit
GET action for each unit
DMA
Put once
Get once
Buffering 1
Buffer once, and put once
Get once, and unbuffer once
2
Buffer for each N L , and put once Get once, and unbuffer for each N L
3
Buffer for each N L , and put once Get once, and unbuffer for each N L
4
Buffer once, and put once
Get once, and unbuffer once
