Three-Dimensional Fluid Code with XcalableMP
167
methods, original source codes are completely same and there are a few directive
differences among three methods.
2.1 Domain Decomposition Methods
The code is parallelized by three different domain decomposition methods, namely
the domain is divided in (a) only Z direction, (b) both Y and Z directions, and (c) all
of X, Y, and Z directions, which are shown in Fig. 1.
Assume LX, LY, and LZ are system mesh sizes of the code in X, Y, and Z
directions, and NX, NY, and NZ are division numbers in X, Y, and Z directions,
respectively. In the case of only Z domain decomposition method, total amount of
communicated data per domain is proportional to LX · LY , which is not depended
on any division numbers, and communication occurs twice per time step. On the
other hand, data proportional to (LX · LY ) ÷ NY must be communicated twice
and that of (LZ · LX) ÷ NZ is also communicated twice per domain per time step
in both Y and Z domain decomposition method. In the case of all of X, Y, and Z
domain decomposition method, data proportional to (LX · LY ) ÷ (NX · NY ), (LZ ·
LX) ÷ (NZ · NX), and (LY · LZ) ÷ (NY · NZ) must be communicated per domain
per time step twice, twice and three times, respectively. Thus total communication
costs in three different domain decomposition methods are depended on a trade-off
between latency and speed. Additional communication is one reduction operation
to obtain a maximum scalar value in whole simulation system.
A node array that corresponds with physical compute units in parallel computer
is defined by node directive, three-dimensional Fortran arrays are decomposed with
template, distribute, and align directives, corresponding DO loops are parallelized
by loop directive and communications between neighboring subdomains are implemented by shadow and reflect directives. The code can be parallelized by using the
“global-view” programming model directives only. Typical XMP Fortran programs
are shown for each decomposition method, in Listing 1 for only Z direction,
Listing 2 for both Y and Z directions, and Listing 3 for all of X, Y, and Z directions.
Fig. 1 IMPACT-3D is parallelized by three different domain decomposition methods. The domain
is divided in (a) only Z direction, (b) both Y and Z directions, and (c) all of X, Y, and Z directions
167
methods, original source codes are completely same and there are a few directive
differences among three methods.
2.1 Domain Decomposition Methods
The code is parallelized by three different domain decomposition methods, namely
the domain is divided in (a) only Z direction, (b) both Y and Z directions, and (c) all
of X, Y, and Z directions, which are shown in Fig. 1.
Assume LX, LY, and LZ are system mesh sizes of the code in X, Y, and Z
directions, and NX, NY, and NZ are division numbers in X, Y, and Z directions,
respectively. In the case of only Z domain decomposition method, total amount of
communicated data per domain is proportional to LX · LY , which is not depended
on any division numbers, and communication occurs twice per time step. On the
other hand, data proportional to (LX · LY ) ÷ NY must be communicated twice
and that of (LZ · LX) ÷ NZ is also communicated twice per domain per time step
in both Y and Z domain decomposition method. In the case of all of X, Y, and Z
domain decomposition method, data proportional to (LX · LY ) ÷ (NX · NY ), (LZ ·
LX) ÷ (NZ · NX), and (LY · LZ) ÷ (NY · NZ) must be communicated per domain
per time step twice, twice and three times, respectively. Thus total communication
costs in three different domain decomposition methods are depended on a trade-off
between latency and speed. Additional communication is one reduction operation
to obtain a maximum scalar value in whole simulation system.
A node array that corresponds with physical compute units in parallel computer
is defined by node directive, three-dimensional Fortran arrays are decomposed with
template, distribute, and align directives, corresponding DO loops are parallelized
by loop directive and communications between neighboring subdomains are implemented by shadow and reflect directives. The code can be parallelized by using the
“global-view” programming model directives only. Typical XMP Fortran programs
are shown for each decomposition method, in Listing 1 for only Z direction,
Listing 2 for both Y and Z directions, and Listing 3 for all of X, Y, and Z directions.
Fig. 1 IMPACT-3D is parallelized by three different domain decomposition methods. The domain
is divided in (a) only Z direction, (b) both Y and Z directions, and (c) all of X, Y, and Z directions
