106
H. Iwashita and M. Nakao
and freeing of MPI windows. In the case of FJ-RDMA, the RS method has no
advantage over the other methods. Since the allocated address is used for registration
to FJ-RDMA, no advantage was found for managing memory outside of the Fortran
system. The unusual connection through the Cray pointer causes degradation of the
Fortran compiler optimization.
3.3 PUT/GET Communication
In order to avoid disturbing the execution on the remote image, PUT and GET communications are always implemented using remote direct memory access (RDMA)
provided by the communication library (except coarrays with pointer/allocatable
structure components). In contrast, local data access is selective between using
direct memory access (DMA) or using a local buffer. For the buffer scheme, one
of the four algorithms will be chosen.
3.3.1 Determining the Possibility of DMA
Coarray variables must be registered when allocated to be the target of RDMA
communication. In contrast, since the local data, which is the source of PUT or
the destination of GET, was not registered or linked to registered information, the
data could not be the target of DMA communication and had to be communicated
via the registered buffer.
When the local data is an entire coarray or a part of coarray, the coarray must
be registered, and efficient DMA-RDMA communication can be made. Since the
analysis at compile time is limited, we implemented the detector in the runtime
library using binary-tree search, as follows.
1. When a chunk of coarray data is registered to the communication library, the
runtime library adds the set of the local address and the size to a sorted table
called SortedChunkTable. The sort key is the local base address of the data.
2. When a chunk of coarray data is deregistered from the communication library,
the runtime library deletes the record in SortedChunkTable.
3. When a PUT or GET runtime library is called corresponding to a reference/definition to a coindexed object/variable, the local address is searched in
SortedChunkTable with binary search. The local data is already registered
if addr i ≤ addr < addr i + size i for any i, where addr is the said local
address and add i and size i are the i-th address and size, respectively, in
SortedChunkTable.
If the communication data is large, then the cost of procedure 3 is relatively small
and is worth using. If the data is small, then the buffering algorithm, as shown in
Sect. 3.3.2, may be better.
H. Iwashita and M. Nakao
and freeing of MPI windows. In the case of FJ-RDMA, the RS method has no
advantage over the other methods. Since the allocated address is used for registration
to FJ-RDMA, no advantage was found for managing memory outside of the Fortran
system. The unusual connection through the Cray pointer causes degradation of the
Fortran compiler optimization.
3.3 PUT/GET Communication
In order to avoid disturbing the execution on the remote image, PUT and GET communications are always implemented using remote direct memory access (RDMA)
provided by the communication library (except coarrays with pointer/allocatable
structure components). In contrast, local data access is selective between using
direct memory access (DMA) or using a local buffer. For the buffer scheme, one
of the four algorithms will be chosen.
3.3.1 Determining the Possibility of DMA
Coarray variables must be registered when allocated to be the target of RDMA
communication. In contrast, since the local data, which is the source of PUT or
the destination of GET, was not registered or linked to registered information, the
data could not be the target of DMA communication and had to be communicated
via the registered buffer.
When the local data is an entire coarray or a part of coarray, the coarray must
be registered, and efficient DMA-RDMA communication can be made. Since the
analysis at compile time is limited, we implemented the detector in the runtime
library using binary-tree search, as follows.
1. When a chunk of coarray data is registered to the communication library, the
runtime library adds the set of the local address and the size to a sorted table
called SortedChunkTable. The sort key is the local base address of the data.
2. When a chunk of coarray data is deregistered from the communication library,
the runtime library deletes the record in SortedChunkTable.
3. When a PUT or GET runtime library is called corresponding to a reference/definition to a coindexed object/variable, the local address is searched in
SortedChunkTable with binary search. The local data is already registered
if addr i ≤ addr < addr i + size i for any i, where addr is the said local
address and add i and size i are the i-th address and size, respectively, in
SortedChunkTable.
If the communication data is large, then the cost of procedure 3 is relatively small
and is worth using. If the data is small, then the buffering algorithm, as shown in
Sect. 3.3.2, may be better.
