88
M. Nakao and H. Murai
1 double L[NB][N];
2 #pragma xmp align L[*][i] with t[*][i]
3 ...
4 int len = N − j − NB;
5 #pragma xmp gmove async (tag)
6 L[0:NB][j+NB:len] = A[j:NB][j+NB:len];
j
A[][]
N
len
NB
L[][]
L[][]
Fig. 11 Panel broadcast in HPL [5]
1 int L_ld, A_ld;
2 xmp_array_lda(xmp_desc_of(L), &L_ld);
3 xmp_array_lda(xmp_desc_of(A), &A_ld);
4 ...
5 cblas_dgemm(..., &L[0][j], L_ld, ..., &A[i][j], A_ld, ...);
Fig. 12 Calling the function cblas_dgemm() in HPL [5]
10
10
10
10
10
10
10
7
6
5
4
3
2
1
Performance (GFlops)
Number of CPUs
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
XMP
MPI
Ratio (XMP/MPI)
1
2
4
2
2
2
6
2
10
2
8
2
12
Performance (GFlops)
10
10
10
10
10
10
10
7
6
5
4
3
2
1
1.2
1.0
0.8
0.6
0.4
0.2
0.0
Perfomance Ratio
XMP
MPI
Ratio (XMP/MPI)
Number of CPUs
1
2
4
2
2
2
6
2
8
The K computer
The COMA system
Fig. 13 Performance results for HPL [5]
cblas_dgemm(). Note that L_ld and A_ld remain unchanged from the beginning of
the program, and so each xmp_array_lda() is called only once.
5.3.3 Evaluation
Figure 13 shows the performance results and performance ratios. XMP’s best
performance results are 402.01 TFlops (76.68% of the peak performance) for
Précédent

- 95/265

Suivant