XcalableMP 2.0 and Future Directions
259
4 Retrospectives and Challenges for Future PGAS Models
Since 2007, we have been developing the XcalableMP PGAS language and its
reference implementation by the Omni compiler.
In this section, the challenges for future PGAS models are presented with some
retrospectives on our project.
4.1 Low-Level Communication Layer for PGAS Model
PGAS is implemented by Remote Memory Access (RMA) providing light-weight
one-sided communication and low overhead synchronization semantics. For programmers, both PGAS and RMA are programming interfaces and offer several
constructs such as remote read/write and synchronizations.
Remote Direct Memory Access (RDMA) is a mechanism (operation) to access
data in remote memory by giving address in (shared) address space. It can be
done without involving the CPU or OS at the destination node. Recent advanced
interconnect such as Cray Aries interconnect and Fujitsu Tofu of K computer and
Tofu-D of Fugaku support remote DMA operations which strongly support efficient
one-sided communication.
For the most PGAS runtimes, one-sided communication operations such as
Remote Direct Memory Access (RDMA) functions in the MPI are used to implement remote put/get operations in the PGAS languages. Although MPI3 provides
several RMA APIs as library interface, the advantages of direct use of RMA/RDMA
Operations are as follows:
• Multiple data transfers can be performed with a single synchronization operation.
• Some irregular communication patterns can be more economically expressed.
• The RDMA can be significantly faster than send/receive on systems with
hardware support for remote memory access.
We found the multiple data transfers for the stencil computation can be optimized
by using a single synchronization operation at the end [13]. As described in Chap. 3,
our XMP Coarray were implemented by both MPI and Fujitsu low-level Tofu API.
In case of MPI, we used “passive target” mode in MPI one-sided communication. It
is noted that the MPI flush operation and synchronization do not sometimes match
to implement “sync_images”, and the complex “window” management to expose
the memory as a coarray. Finally, Fujitsu RDMA interface is much faster than MPI
in the K computer.
Other problem is the communication in the multithreaded environment.
As described in the previous sections, we found the performance problem of
MPI_THREAD_MULTIPLE. As the connection-less semantics of RDMA would
be suited to communications in multithreaded environment, we believe that a new
design of low-level communication layer would be a desirable solution in near
future.
Précédent

- 262/265

Suivant