11.5 Parallel Computing in CFD
365
where NCv is the total number of CVs, T is the time per floating point operation and is is the number of floating point operations per CV required
to reach convergence. For a parallel algorithm executed on n processors, the
total execution time consists of computing and communication time:
T - ~ i d C
n -
+ T z m = N F T ~ , + T F m ,
(11.20)
where N F is the number of CVs in the largest sub-domain and Tiom is
the total communication time during which calculation cannot take place.
Inserting these expressions into the definition of the total efficiency yields:
tot - - 2 - -
NCV7is
- -
En
n Tn n (N?rin + T,Com)
This equation is not exact, since the number of floating point operations
per CV is not constant (due to branching in the algorithm and the fact
that boundary conditions affect only some CVs). However, it is adequate to
identify the major factors affecting the total efficiency. The meanings of these
factors are:
EiUm - The numerical efficiency accounts for the effect of the change in
the number of operations per grid node required to reach convergence due
to modification of the algorithm to allow parallelization;
E r - The parallel efficiency accounts for the time spent on communication during which computation cannot take place;
EAb - The load balancing efficiency accounts for the effect of some processors being idle due to uneven load.
When the parallelization is performed in both time and space, the overall
efficiency is equal to the product of time and space efficiencies.
The total efficiency is easily determined by measuring the computing time
necessary to obtain the converged solution. The parallel efficiency cannot be
measured directly, since the number of inner iterations is not the same for
all outer iterations (unless it is fixed by the user). However, if we execute
a certain number of outer iterations with a fixed number of inner iterations
per outer iteration on 1 and n processors, the numerical efficiency is unity
and the total efficiency is then the product of the parallel and load balancing
efficiencies. If the load balancing efficiency is reduced to unity by making all
sub-domains equal, we obtain the parallel efficiency. Some computers have
tools which allow operation counts to be performed; then the numerical efficiency can be directly measured.
For both space and time domain decomposition, all three efficiencies are
usually reduced as the number of processors is increased for a given grid. This
decrease is both non-linear and problem-dependent. The parallel efficiency is
especially affected, since the time for LC is almost constant, the time for GC
Précédent

- 375/779

Suivant