Implementation and Performance Evaluation of Omni Compiler
85
problem size to be equal to the minimum size. As for coarray syntax, Omni compiler
uses FJRDMA on the K computer and uses GASNet on the COMA system.
5.2 EP STREAM Triad
5.2.1 Design
STREAM measures the memory bandwidth to use simple vector kernel (a ← b +
αc). STREAM is so straightforward that its kernel does not require communication.
5.2.2 Implementation
Figure 7 shows a part of the STREAM code. In line 1, the node directive declares a
node array p to parallelize the program. In line 2, normal arrays a[], b[], and c[], and
a scalar value scalar are declared. In lines 5 and 14, the barrier directive is inserted
before xmp_wtime() to measure time. The directives of lines 8–9 are optimization
directives for the Fujitsu compiler. While the #pragma loop xfill ensures one cache
line to store write-only data, the #pragma loop noalias indicates that different
pointer variables cannot possibly indicate the same storage area. These optimization
directives are used on only the K computer. In lines 10–12, STREAM kernel is
parallelized by the OpenMP parallel directive. In line 17, local_performance()
1 #pragma xmp nodes p[*]
2 double a[N], b[N], c[N], scalar;
3 ...
4 for(k=0;k
5 #pragma xmp barrier
6
times[k] = −xmp_wtime();
7
8 #pragma loop xfill
9 #pragma loop noalias
10 #pragma omp parallel for
11
for (i=0; i
12
a[i] = b[i] + scalar * c[i];
13
14 #pragma xmp barrier
15
times[k] += xmp_wtime();
16 }
17 double performance = local_performance(time, TIMES, N);
18 #pragma xmp reduction(+:performance)
Fig. 7 Part of the STREAM code [5]
85
problem size to be equal to the minimum size. As for coarray syntax, Omni compiler
uses FJRDMA on the K computer and uses GASNet on the COMA system.
5.2 EP STREAM Triad
5.2.1 Design
STREAM measures the memory bandwidth to use simple vector kernel (a ← b +
αc). STREAM is so straightforward that its kernel does not require communication.
5.2.2 Implementation
Figure 7 shows a part of the STREAM code. In line 1, the node directive declares a
node array p to parallelize the program. In line 2, normal arrays a[], b[], and c[], and
a scalar value scalar are declared. In lines 5 and 14, the barrier directive is inserted
before xmp_wtime() to measure time. The directives of lines 8–9 are optimization
directives for the Fujitsu compiler. While the #pragma loop xfill ensures one cache
line to store write-only data, the #pragma loop noalias indicates that different
pointer variables cannot possibly indicate the same storage area. These optimization
directives are used on only the K computer. In lines 10–12, STREAM kernel is
parallelized by the OpenMP parallel directive. In line 17, local_performance()
1 #pragma xmp nodes p[*]
2 double a[N], b[N], c[N], scalar;
3 ...
4 for(k=0;k
6
times[k] = −xmp_wtime();
7
8 #pragma loop xfill
9 #pragma loop noalias
10 #pragma omp parallel for
11
for (i=0; i
a[i] = b[i] + scalar * c[i];
13
14 #pragma xmp barrier
15
times[k] += xmp_wtime();
16 }
17 double performance = local_performance(time, TIMES, N);
18 #pragma xmp reduction(+:performance)
Fig. 7 Part of the STREAM code [5]
