Implementation and Performance Evaluation of Omni Compiler
87
K computer, and 11.55 TB/s for 128 compute nodes on the COMA system. The
values of the performance ratio are between 0.99 and 1.00 on both systems.
5.3 High-Performance Linpack
5.3.1 Design
HPL evaluates the floating point rate of execution for solving a linear system of
equations. The performance result has been used in the TOP500 list (https://www.
top500.org). To achieve a good load balance on HPL, we distribute the main array
in a block-cyclic manner. Moreover, in order to achieve high performance with
portability, our implementation calls BLAS [6] to perform the matrix operations.
These techniques are inherited from the MPI implementation.
5.3.2 Implementation
Figure 10 shows that each dimension of the coefficient matrix A[][] is distributed
in the block-cyclic manner. The template and the nodes directives declare a twodimensional template t and node array p. The distribute directive distributes t onto
Q × P nodes with the same block size NB. The align directive aligns A[][] with t.
HPL has an operation in which a part of the coefficient matrix is broadcast to the
other process columns asynchronously. This operation, called “panel broadcast,”
is one of the most important operations for overlapping panel factorizations and
data transfer. Figure 11 shows the implementation that uses the gmove directive
with the async clause. The second dimension of array L[][] is also distributed in a
block-cyclic manner and L[][] is replicated. Thus, the gmove directive broadcasts
elements A[j:NB][j+NB:len] to L[0:NB][j+NB:len] asynchronously.
Figure 12 shows that cblas_dgemm(), which is a BLAS function for a
matrix multiplication, applies the distributed arrays L[][] and A[][]. Note
that cblas_dgemm() is executed by multiple threads locally. In lines 2–3,
xmp_desc_of() gets descriptors of L[][] and A[][], and xmp_array_lda() gets
the leading dimensions L_ld and A_ld. In line 5, the L_ld and A_ld are used in
Fig. 10 Block-cyclic
distribution in HPL [5]
t
p[0][0]
p[1][0]
p[0][1]
p[1][1]
N
NB
j
i
87
K computer, and 11.55 TB/s for 128 compute nodes on the COMA system. The
values of the performance ratio are between 0.99 and 1.00 on both systems.
5.3 High-Performance Linpack
5.3.1 Design
HPL evaluates the floating point rate of execution for solving a linear system of
equations. The performance result has been used in the TOP500 list (https://www.
top500.org). To achieve a good load balance on HPL, we distribute the main array
in a block-cyclic manner. Moreover, in order to achieve high performance with
portability, our implementation calls BLAS [6] to perform the matrix operations.
These techniques are inherited from the MPI implementation.
5.3.2 Implementation
Figure 10 shows that each dimension of the coefficient matrix A[][] is distributed
in the block-cyclic manner. The template and the nodes directives declare a twodimensional template t and node array p. The distribute directive distributes t onto
Q × P nodes with the same block size NB. The align directive aligns A[][] with t.
HPL has an operation in which a part of the coefficient matrix is broadcast to the
other process columns asynchronously. This operation, called “panel broadcast,”
is one of the most important operations for overlapping panel factorizations and
data transfer. Figure 11 shows the implementation that uses the gmove directive
with the async clause. The second dimension of array L[][] is also distributed in a
block-cyclic manner and L[][] is replicated. Thus, the gmove directive broadcasts
elements A[j:NB][j+NB:len] to L[0:NB][j+NB:len] asynchronously.
Figure 12 shows that cblas_dgemm(), which is a BLAS function for a
matrix multiplication, applies the distributed arrays L[][] and A[][]. Note
that cblas_dgemm() is executed by multiple threads locally. In lines 2–3,
xmp_desc_of() gets descriptors of L[][] and A[][], and xmp_array_lda() gets
the leading dimensions L_ld and A_ld. In line 5, the L_ld and A_ld are used in
Fig. 10 Block-cyclic
distribution in HPL [5]
t
p[0][0]
p[1][0]
p[0][1]
p[1][1]
N
NB
j
i
