364
11. Efficiency and Accuracy Improvement
Fig. 11.14. Structure of
the global coefficient matrix
when four time time steps
are calculated in parallel
to multilevel schemes; in that case the processors have to send and receive
data from more than one time level.
11.5.4 Efficiency of Parallel Computing
The analysis of the performance of parallel programs is usually measured by
the speed-up factor and efficiency defined by:
Here T, is the execution time for the best serial algorithm on a single processor
and T, is the execution time for the parallelized algorithm using n processors.
In general T, # TI, as the best serial algorithm may be different from the
best parallel algorithm; one should not base the efficiency on the performance
of the parallel algorithm executed on a single processor.
The speed-up is usually less than n (the ideal value), so the efficiency
is usually less than 1 (or 100%). However, when solving coupled non-linear
equations, it may turn out that solution on two or four processors is more
efficient than on 1 processor so, in principle, efficiencies higher than 100% are
possible (the increase is often due to the better use of cash memory when a
smaller problem is solved by one processor).
Although not necessary, the processors are usually synchronized a t the
start of each iteration. Since the duration of one iteration is dictated by
the processor with the largest number of CVs, other processors experience
some idle time. Delays may also be due to different boundary conditions in
different sub-domains, different numbers of neighbors, or more complicated
communication.
The computing time Ts may be expressed as:
Précédent

- 374/431

Suivant