XcalableMP 2.0 and Future Directions
255
1 double A[nt][nt][ts*ts], B[ts*ts], C[nt][ts*ts];
2 #pragma xmp nodes P(*)
3 #pragma xmp template T(0:nt−1)
4 #pragma xmp distribute T(cyclic) onto P
5 #pragma xmp align A[*][i][*] with T(i)
6
7 #pragma xmp tasklets
8 for (int k = 0; k < nt; k++) {
9 #pragma xmp tasklet out(A[k][k]) on T(k)
10
potrf(A[k][k]);
11
12 #pragma xmp tasklet gmove in(A[k][k]) out(B) on T(k:)
13
B[:] = A[k][k][:];
14
15
for (int i = k + 1; i < nt; i++) {
16 #pragma xmp tasklet in(B) out(A[k][i]) on T(i)
17
trsm(B, A[k][i]);
18
19 #pragma xmp tasklet gmove in(A[k][i]) out(C[i]) on T(i:)
20
C[i][:] = A[k][i][:];
21
}
22
for (int i = k + 1; i < nt; i++) {
23
for (int j = k + 1; j < i; j++) {
24 #pragma xmp tasklet in(A[k][i], C[j]) out(A[j][i]) on T(j)
25
gemm(A[k][i], C[j], A[j][i]);
26
}
27 #pragma xmp tasklet in(A[k][i]) out(A[i][i]) on T(i)
28
syrk(A[k][i], A[i][i]);
29
}
30 }
Fig. 7 Blocked Cholesky factorization code in the XMP tasklet directive
We also show the result of these benchmarks implemented by MPI and OmpSs as
“Task (OmpSs).” This implementation is described in the in, out, and inout
clauses with the OmpSs task directive. The parallel and single regions are
not required in the OmpSs programming model. Except for these differences, this is
almost the same as the Task version.
We evaluated these benchmarks on the following node configurations. For the
Oakforest-PACS system, it is on 1–32 nodes, one process per node, 64 cores per
process, and one thread per core. The problem size of these benchmarks is set
by a matrix size of 32,768 × 32,768 and a block size of 512 × 512 in double
precision arithmetic. The matrix is distributed by a two-dimensional block-cyclic
data distribution in blocked Cholesky factorization.
Figure 8 illustrates the performance and breakdown of blocked Cholesky factorization on the Oakforest-PACS. The breakdown indicates the average time required
Précédent

- 258/265

Suivant