Implementation and Performance Evaluation of Omni Compiler
95
MPI implementation for 2048 Bytes or greater transfer size on the K computer.
In contrast, the results on the COMA system show that the latency of the XMP
implementation is always worse than that of the MPI implementation on the COMA
system. The latency of XMP with FJRDMA at 4096 Bytes, which is the average
data chunk size, is 5.83 μs and the latency of MPI is 6.89 μs on the K computer.
The latency of XMP with GASNet is 5.05 μs and that of MPI is 3.37 μs on the
COMA system. Thus, we consider the reason for the performance difference of
RandomAccess is the communication performance. The performance difference
is also due to the differences in the synchronization mechanism of the one-sided
XMP coarray and the two-sided MPI functions. Note that a real application would
not synchronize after every one-sided communication. It is expected that a single
synchronization should occur after multiple one-sided communications to achieve
higher performance.
In addition, the performance results for HPL and FFT are slightly different
from those for the MPI implementations. We consider that these differences are
caused by small differences in the implementations. In HPL, for the panel-broadcast
operation, the XMP implementation uses the gmove directive with the async clause,
which calls MPI_Ibcast() internally. In contrast, the MPI implementation uses
MPI_Send() and MPI_Recv() to perform the operation by the “modified increasing
ring” [11]. In FFT, the XMP implementation uses XMP in Fortran, but the MPI
implementation uses C language. Both implementations call the same FFTE library.
In addition, the MPI implementation uses MPI_Alltoall() to transpose a matrix.
Since xmp_transpose() calls MPI_Alltoall() internally, the performance levels
for both xmp_transpose() and MPI_Alltoall() must be the same. Therefore, the
language differences and refactoring may have caused the performance difference.
6 Conclusion
The chapter describes the implementation and performance evaluation of Omni
compiler. We evaluate the performance of the HPCC benchmark in XMP on the
K computer up to 16,384 compute nodes and a generic cluster system up to
128 compute nodes. The performance results for the XMP implementations are
almost the same as those for the MPI implementations in many cases. Moreover,
it demonstrates that the global-view and the local-view memory models are useful
to develop the HPCC benchmark.
References
1. Programming Environment Research Team, https://pro-env.riken.jp
2. High Performance Computing System laboratory, University of Tsukuba, Japan, https://www.
hpcs.cs.tsukuba.ac.jp
95
MPI implementation for 2048 Bytes or greater transfer size on the K computer.
In contrast, the results on the COMA system show that the latency of the XMP
implementation is always worse than that of the MPI implementation on the COMA
system. The latency of XMP with FJRDMA at 4096 Bytes, which is the average
data chunk size, is 5.83 μs and the latency of MPI is 6.89 μs on the K computer.
The latency of XMP with GASNet is 5.05 μs and that of MPI is 3.37 μs on the
COMA system. Thus, we consider the reason for the performance difference of
RandomAccess is the communication performance. The performance difference
is also due to the differences in the synchronization mechanism of the one-sided
XMP coarray and the two-sided MPI functions. Note that a real application would
not synchronize after every one-sided communication. It is expected that a single
synchronization should occur after multiple one-sided communications to achieve
higher performance.
In addition, the performance results for HPL and FFT are slightly different
from those for the MPI implementations. We consider that these differences are
caused by small differences in the implementations. In HPL, for the panel-broadcast
operation, the XMP implementation uses the gmove directive with the async clause,
which calls MPI_Ibcast() internally. In contrast, the MPI implementation uses
MPI_Send() and MPI_Recv() to perform the operation by the “modified increasing
ring” [11]. In FFT, the XMP implementation uses XMP in Fortran, but the MPI
implementation uses C language. Both implementations call the same FFTE library.
In addition, the MPI implementation uses MPI_Alltoall() to transpose a matrix.
Since xmp_transpose() calls MPI_Alltoall() internally, the performance levels
for both xmp_transpose() and MPI_Alltoall() must be the same. Therefore, the
language differences and refactoring may have caused the performance difference.
6 Conclusion
The chapter describes the implementation and performance evaluation of Omni
compiler. We evaluate the performance of the HPCC benchmark in XMP on the
K computer up to 16,384 compute nodes and a generic cluster system up to
128 compute nodes. The performance results for the XMP implementations are
almost the same as those for the MPI implementations in many cases. Moreover,
it demonstrates that the global-view and the local-view memory models are useful
to develop the HPCC benchmark.
References
1. Programming Environment Research Team, https://pro-env.riken.jp
2. High Performance Computing System laboratory, University of Tsukuba, Japan, https://www.
hpcs.cs.tsukuba.ac.jp
