254
M. Sato et al.
preliminary evaluation, we have made a hand-translated MPI and OpenMP code by
using the proposed transformation.
The tasklets directive is converted into the OpenMP parallel and
single directives. The execution node is determined by the on clause, which is
translated to an if statement. The tasklet gmove and tasklet reflect
directives are converted into MPI_Send/Recv(), and these MPI functions are
executed in OpenMP tasks with data dependency specified by users. In the case that
an MPI blocking call, such as MPI_Send/Recv(), occurs in these codes, a deadlock
may occur depending on the task scheduling mechanism, from the combination
of MPI and OpenMP. To prevent this deadlock, in the actual implementation we
used MPI asynchronous communications, such as MPI_Isend/Irecv(), MPI_Test(),
and the OpenMP taskyield directive, which makes the current task become
suspended at the time point at which it is invoked, and may result in switching
to different tasks.
3.4 Preliminary Performance
We measured the performance on the Oakforest-PACS [11] systems at at the
Joint Center for Advanced High-Performance Computing (JCAHPC) [9], under
cooperation with the Center for Computational Sciences, University of Tsukuba and
the Information Technology Center, the University of Tokyo. This system has 8,208
computing nodes, each of which consists of an Intel Xeon Phi (KNL) processor
and the Intel Omni-Path architecture as an interconnection. In this evaluation, we
selected the Flat and Quadrant modes for KNL. While the Intel Xeon Phi 7250
has 68 cores, a 64 core usage per node is recommended in this system. Some
cores are used to assist the OS, interrupt handling, and for communication progress.
Moreover, in order to avoid OS jitters, only core 0 is set to receive OS interruptions.
We used blocked Cholesky factorization as our benchmark. It calculates the
decomposition of a Hermitian positive-definite blocked matrix into the product of
a lower triangular matrix and its conjugate transpose. The calculation consists of
four BLAS or LAPACK functions, POTRF, TRSM, GEMM, and SYRK, which are
performed in block units. Figure 7 shows the Blocked Cholesky factorization code
in the XMP tasklet directive.
We compare the performance in two parallelization approaches, “Parallel Loop”
and “Task,” in MPI and OpenMP. The “Parallel Loop” version is the conventional
barrier-based implementation, described by work sharing for loops using the
parallel for directive and independent tasks using the task directive without
the depend clause. Although this version of blocked Cholesky factorization is
applied on the overlap of the communication and computation at the process level, it
performs the global synchronization in work sharing. The “Parallel Loop” version of
the Laplace equation solver does not include the overlap of the communication and
computation. The “Task” version is implemented using our proposed model, based
on task dependency using the depend clause, instead of global synchronization.
Précédent

- 257/265

Suivant