196
K. Tsugane et al.
Fig. 15 Breakdown of the minimum and maximum calculation times on the radial domain
decomposition using 512 processes
are calculated based on the formula used to describe the torus shape, which implies
that an error is incurred during integer rounding to determine the number of grids.
Table 5 shows that the number of total grid points assigned to the processes with the
maximum and minimum calculation times differed greatly. Also, Fig. 15 shows the
breakdown of the minimum and maximum calculation times on the radial domain
decomposition 512 processes. The difference between the calculation times of the
grid-related works, such as charge, push, poisson, field, and smooth,
increased on radial domain decomposition. During each time step, the computation
of all processes must be bounded as a barrier operation and the increase in the integer
rounding error according to the problem size (i.e., weak scaling) causes a greater
load imbalance, which degrades the overall performance.
On the other hand, the communication time of XMP-localview and XMPhybridview on the radial dimension increases as the number of nodes increased
compared with the MPI, as shown in Fig. 13. We explored the number of send calls
and each communication size because the performance of communication on XMP
and MPI are reversed at about 65,536 Bytes according to Fig. 11. Figure 16 shows
the number of send calls in process number 0 on each domain decomposition
classified as the communication size of more than 65,536 Bytes and 65,536 Bytes
or less. In the radial domain decomposition, the number of send calls at more
than 65,536 Bytes increases compared with the toroidal and particle decomposition.
Therefore, the performance of XMP-localview and XMP-hybridview is degraded
compared with MPI. The results were the XMP-localview implementation obtains
approximately the same performance as the MPI implementation while the performance degradation using XMP-hybridview is increased by up to 20% compared
with the MPI implementation.
With strong scaling, Figs. 17 and 18 show the elapsed time for both calculation
and communication of MPI, XMP-localview, and XMP-hybridview, where the
decomposition on the radial and particle dimensions, respectively. The performances of XMP-localview and XMP-hybridview on the particle dimension are
Précédent

- 201/265

Suivant