Implementation and Performance Evaluation of Omni Compiler
93
XMP
MPI
Ratio (XMP/MPI)
Number of CPUs
10
10
10
10
10
10
10
3
2
1
0
-1
-2
-3
Performance (GUPS)
1
2
4
2
2
2
6
2
10
2
8
2
12 2
14
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
XMP
MPI
Ratio (XMP/MPI)
10
10
10
10
10
10
10
3
2
1
0
-1
-2
-3
Performance (GUPS)
1
2
4
2
2
2
6
2
8
Number of CPUs
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
The K computer
The COMA system
Fig. 18 Performance results for RandomAccess [5]
5.6 Discussion
We implement STREAM, HPL, and FFT using the global-view memory model,
which enables programmers to develop the parallel codes from the sequential
codes using the XMP directives and functions easily. Specifically, in order to
implement the parallel STREAM code, a programmer only adds the XMP directives
into the sequential STREAM code. The XMP directives and existing directives,
such as OpenMP directives and Fujitsu directives, can coexist. Moreover, existing
high-performance libraries, such as BLAS and FFTE, can be used with an XMP
distributed array. These features improve the portability and performance of XMP
applications.
We also implement RandomAccess using the local-view memory model, where
the coarray syntax enables a programmer to transfer data intuitively. In the
evaluation, the performance of the XMP implementation is better than that of
the MPI implementation on the K computer, but is worse than that of the MPI
implementation on the COMA system.
To clarify the reason why XMP performance is dropped on the COMA system,
we develop a simple ping-pong benchmark using the local-view memory model.
The benchmark measures the latency for transferring data repeatedly between two
nodes. For comparison purposes, we also implement one using MPI_Isend() and
MPI_Irecv() that are used in the MPI version RandomAccess.
Figure 19 shows parts of the codes. In XMP of Fig. 19, in line 5, p[0] puts a
part of src_buf[] into dst_buf[] in p[1]. In line 6, the post directive ensures the
completion of the coarray operation of line 5 and sends a notification to p[1]. In
line 10, p[1] waits until receiving the notification from p[0]. In line 11, p[1] puts
a part of src_buf[] into dst_buf[] in p[0]. In line 12, the post directive ensures the
completion of the coarray operation of line 11 and sends a notification to p[0]. In
line 7, p[0] waits until receiving the notification from p[1]. Figure 19 also shows
the ping-pong benchmark that uses MPI functions.
Figure 20 shows the latency for transferring data. The results on the K computer
show that the latency for the XMP implementation is better than that for the
Précédent

- 100/265

Suivant