86
M. Nakao and H. Murai
calculates the performance on each node locally. In line 18, the reduction directive
performs a reduction operation among nodes to calculate the total performance.
5.2.3 Evaluation
First of all, in order to consider the effectiveness of #pragma loop xfill and
#pragma loop noalias, we evaluate STREAM with and without these directives
on a single node of the K computer. We also insert these directives into the
MPI implementation for evaluation. Figure 8 shows that the performance results
with these directives are about 1.46 times better than those without the directives.
Therefore, we use the directives in next evaluations.
Figure 9 shows the performance results and a comparative performance evaluation of both implementations. The comparative performance evaluation is called the
“performance ratio.” When the performance ratio is greater than 1, the performance
result of the XMP implementation is better than that of the MPI implementation.
XMP’s best performance results are 706.38 TB/s for 16,384 compute nodes on the
Performance (GB/s)
50
40
30
20
10
0
without directives
MPI
XMP
with directives
Fig. 8 Preliminary evaluation of STREAM [5]
Number of CPUs
10
10
10
10
10
10
10
7
6
5
4
3
2
1
Performance (GB/s)
1
2
4
2
2
2
6
2
10
2
8
2
12 2
14
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
XMP
MPI
Ratio (XMP/MPI)
10
10
10
10
10
10
10
7
6
5
4
3
2
1
Performance (GB/s)
1
2
4
2
2
2
6
2
8
Number of CPUs
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
XMP
MPI
Ratio (XMP/MPI)
The K computer
The COMA system
Fig. 9 Performance results for STREAM [5]
Précédent

- 93/265

Suivant