11.5 Parallel Computing in CFD
361
The SIP solver is easily adapted to this method. Each diagonal block
matrix Mii is decomposed into L and U matrices in the normal way; the global
iteration matrix M = LU is not the one found in the single processor case.
After one iteration is performed on each sub-domain, one has to exchange the
updated values of the unknown q5m so that the residual pm can be calculated
at nodes near sub-domain boundaries.
When SIP solver is parallelized in this way, the performance deteriorates
as the number of processors becomes large; the number of iterations may
double when the number of processors is increased from one to 100. However,
if the inner iterations do not have to be converged very tightly, as is the
case for some implicit schemes described in Chap. 7, parallel SIP can be
quite efficient because SIP tends to reduce the error rapidly in the first few
iterations. Especially if the multigrid method is used to speed-up the outer
iterations, the total efficiency is quite high (80% to 90%; see Schreck and
PeriC, 1993, and Lilek et al., 1995, for examples). A 2D flow prediction code
parallelized in this way is available via Internet; see appendix for details.
1=1 -*---20 t
Coupled preconditioner
1=2 +
Convergence criterion = 0.0 1
1=3 ...+ .....,,.
0 1
0
10
20
30
40
50
60
No. of proc.
Fig. 11.13. Number of iterations in the ICCG solver as a function of the number of
processors (uniform grid with 643 CVs, Poisson equation with Neumann boundary
conditions, LC after each pre-conditioner sweep, 1 sweeps per CG iteration, residual
norm reduced two orders of magnitude; from Seidl, 1997)
Conjugate gradient based methods can also be parallelized using the above
approach. Below we present a pseudo-code for the preconditioned CG solver.
It was found (Seidl et al., 1995) that the best performance is achieved by
performing two pre-conditioner sweeps per CG iteration, on either single- or
multi-processors. Results of solving a Poisson equation with Neumann bound-
361
The SIP solver is easily adapted to this method. Each diagonal block
matrix Mii is decomposed into L and U matrices in the normal way; the global
iteration matrix M = LU is not the one found in the single processor case.
After one iteration is performed on each sub-domain, one has to exchange the
updated values of the unknown q5m so that the residual pm can be calculated
at nodes near sub-domain boundaries.
When SIP solver is parallelized in this way, the performance deteriorates
as the number of processors becomes large; the number of iterations may
double when the number of processors is increased from one to 100. However,
if the inner iterations do not have to be converged very tightly, as is the
case for some implicit schemes described in Chap. 7, parallel SIP can be
quite efficient because SIP tends to reduce the error rapidly in the first few
iterations. Especially if the multigrid method is used to speed-up the outer
iterations, the total efficiency is quite high (80% to 90%; see Schreck and
PeriC, 1993, and Lilek et al., 1995, for examples). A 2D flow prediction code
parallelized in this way is available via Internet; see appendix for details.
1=1 -*---20 t
Coupled preconditioner
1=2 +
Convergence criterion = 0.0 1
1=3 ...+ .....,,.
0 1
0
10
20
30
40
50
60
No. of proc.
Fig. 11.13. Number of iterations in the ICCG solver as a function of the number of
processors (uniform grid with 643 CVs, Poisson equation with Neumann boundary
conditions, LC after each pre-conditioner sweep, 1 sweeps per CG iteration, residual
norm reduced two orders of magnitude; from Seidl, 1997)
Conjugate gradient based methods can also be parallelized using the above
approach. Below we present a pseudo-code for the preconditioned CG solver.
It was found (Seidl et al., 1995) that the best performance is achieved by
performing two pre-conditioner sweeps per CG iteration, on either single- or
multi-processors. Results of solving a Poisson equation with Neumann bound-