11.5 Parallel Computing in CFD
367
Table 11.2. Numerical efficiency for various domain decompositions in space and
time for the calculation of cavity flow with an oscillating lid
Decomposition in Mean number of Mean number of Numerical
space and time
outer iterations
inner iterations
efficiency
x x y x z x t
per time step
per time step
(in %)
1 x 1 x 1 x 1
11.3
359
100
1 x 2 ~ 2 ~ 1 11.6
417
90
1 x 4 ~ 4 ~ 1 11.3
635
68
the sub-domain. With time parallelism, one can assemble the new coefficient
and source matrices while LC is taking place. Even the GC in a conjugate
gradient solver, which appears to hinder execution, can be overlapped with
computation if the algorithm is rearranged as suggested by Demmel et al.
(1993). Convergence checking can be skipped in the early stages of computation, or the convergence criterion can be rearranged to monitor the residual
level at the previous iteration, and one can base the decision to stop on that
value and the rate of reduction.
PeriC and Schreck (1995) analyzed the possibilities of overlapping communication and computation in more detail and found that it can significantly
improve parallel efficiency. New hard- and software are likely t o allow concurrency of computation and communication, so one can expect that parallel
efficiency can be optimized. One of the main concerns for the developers of
parallel implicit CFD algorithms is numerical efficiency. It is essential that
the parallel algorithm not need many more computing operations than the
serial algorithm for the same accuracy. Results show that parallel computing
can be efficiently used in CFD. The use of workstation clusters is especially
useful with this respect, as they are available to almost all users and big problems are not solved all the time. It is expected that most computers (PCs,
workstations and mainframes) will be multiprocessor machines in the future;
it is therefore essential to have parallel processing in mind when developing
new solution methods.
367
Table 11.2. Numerical efficiency for various domain decompositions in space and
time for the calculation of cavity flow with an oscillating lid
Decomposition in Mean number of Mean number of Numerical
space and time
outer iterations
inner iterations
efficiency
x x y x z x t
per time step
per time step
(in %)
1 x 1 x 1 x 1
11.3
359
100
1 x 2 ~ 2 ~ 1 11.6
417
90
1 x 4 ~ 4 ~ 1 11.3
635
68
the sub-domain. With time parallelism, one can assemble the new coefficient
and source matrices while LC is taking place. Even the GC in a conjugate
gradient solver, which appears to hinder execution, can be overlapped with
computation if the algorithm is rearranged as suggested by Demmel et al.
(1993). Convergence checking can be skipped in the early stages of computation, or the convergence criterion can be rearranged to monitor the residual
level at the previous iteration, and one can base the decision to stop on that
value and the rate of reduction.
PeriC and Schreck (1995) analyzed the possibilities of overlapping communication and computation in more detail and found that it can significantly
improve parallel efficiency. New hard- and software are likely t o allow concurrency of computation and communication, so one can expect that parallel
efficiency can be optimized. One of the main concerns for the developers of
parallel implicit CFD algorithms is numerical efficiency. It is essential that
the parallel algorithm not need many more computing operations than the
serial algorithm for the same accuracy. Results show that parallel computing
can be efficiently used in CFD. The use of workstation clusters is especially
useful with this respect, as they are available to almost all users and big problems are not solved all the time. It is expected that most computers (PCs,
workstations and mainframes) will be multiprocessor machines in the future;
it is therefore essential to have parallel processing in mind when developing
new solution methods.
