232
M. Tsuji et.al.
instances computes an Arnoldi iteration asynchronously over n nodes and sends the
resulting (m, H, V , f ) to the data server. The data server keeps the best result and
sends it to each IRAM. Each IRAM restarts with the (m, H, V , f ) sent by the data
server.
6.2 Experiments
Here, we show the result of experiments on the T2K-Tsukuba supercomputer. The
specification of the T2K Tsukuba is shown in Table 4. In the experiments, we use a
matrix called Schenk/nlpkkt240 from the SuiteSparse Matrix Collection [11], where
n = 27,993,600 and the number of non-zero elements are 760,648,352.
Figure 13 shows the results of MIRAM with IRAM solvers of m = 24, 32, 40
(left) and 3 independent runs of IRAM solvers of the m = 24, 32, 40 (right). While
MIRAM converged around 450 iterations, none of 3 IRAMs could not converge
until 500 iterations.
This MIRAM example shows that by using the mSPMD programming model,
two different accelerations can be achieved. While the workflow programming
model of the mSPMD accelerates the convergence of the Arnoldi iterations, the
distributed parallel programming model speeds up each iteration of the Arnoldi
method.
Table 4 Specification of T2K Tsukuba
CPU
Opteron Barcelona B8000, 4 cores, 4 sockets, 2.3 GHz
Memory
32 GB
Network
Fat-tree, full-bisection interconnection quad-rail of InfiniBand
Fig. 13 Results of MIRAM with 3 IRAM solvers (right) and 3 independent runs of IRAM (left)
Précédent

- 236/265

Suivant