366
11. Efficiency and Accuracy Improvement
increases, and the computing time per processor decreases due t o reduced subdomain size. For time parallelization, the time for GC increases while the LC
and computing times remain the same when more time steps are calculated
in parallel for the same problem size. However, the numerical efficiency will
decrease disproportionately if the number of processors is increased beyond a
certain limit (which depends on the number of outer iterations per time step).
Optimization of the load balancing is difficult in general, especially if the grid
is unstructured and local refinement is employed. There are algorithms for
optimization, but they may take more time than the flow computation!
Parallel efficiency can be expressed as a function of three main parameters:
0 set-up time for data transfer (called latency time);
0 data transfer rate (usually expressed in Mbytesls);
0 computing time per floating point operation (usually expressed in Mflops).
For a given algorithm and communication pattern, one can create a model
equation to express the parallel efficiency as a function of these parameters
and the domain topology. Schreck and PeriC (1993) presented such a model
and showed that the parallel efficiency can be fairly well predicted. One can
also model the numerical efficiency as a function of alternatives in the solution algorithm, the choice of solver and the coupling of the sub-domains.
However, empirical input based on experience with similar flow problems is
necessary, since the behavior of the algorithm is problem-dependent. These
models are useful if the solution algorithm allows alternative communication patterns; one can choose the one most suitable for the computer used.
For example, one can exchange data after each inner iteration, after every
second inner iteration, or only after each outer iteration. One can employ
one, two, or more pre-conditioner iterations per conjugate gradient iteration;
the pre-conditioner iterations may include local communication after each
step or only a t the end. These options affect both the numerical and parallel
efficiency; a trade-off is necessary to find an optimum.
Combined space and time parallelization is more efficient than pure spatial parallelization because, for a given problem size, the efficiency goes down
as the number of processors increases. Table 11.5.4 shows results of the computation of the unsteady 3D flow in a cubic cavity with an oscillating lid a t a
maximum Reynolds number of lo4, using a 32 x 32 x 32 CV grid and a time
step of At = T/200, where T is the period of the lid oscillation. When sixteen
processors were used with four time steps calculated in parallel and the space
domain decomposed into four sub-domains, the total numerical efficiency was
97%. If all processors are used solely for spatial or temporal decomposition,
the numerical efficiency drops below 70%.
Communication between processors halts computation on many machines.
However, if the communication and computation could take place simultaneously (which is possible on some new parallel computers), many parts of the
solution algorithm could be rearranged to take advantage of this. For example,
while LC takes place in the solver, one can do computation in the interior of
Précédent

- 376/779

Suivant