Multi-SPMD Programming Model with YML and XcalableMP
235
# of blocks 1 × 1
2 × 2
4 × 4 8 × 8
Block size 20,480 2 10,240 2 5120 2 2560 2
1024 cores (64 nodes) are used for each workflow, and 64–1024 cores are
assigned for each task in a workflow.
Firstly, we have considered the overhead of the heartbeat messages used to
detect errors in remote programs. Figure 15 shows the performance of the normal
and fault-tolerant mSPMD programming executions using between 64 and 1024
compute cores per task. The dotted lines are the results of fault-tolerant mSPMD
programming executions, and the solid lines are those of the regular mSPMD
programming executions. As shown in the figure, the best combination of the
number of blocks and the number of processes per task is 4 × 4 blocks and 512
processes for both cases of with and without fault-tolerance support. The overhead
of using a heartbeat message is very small and is 2.3% on average and 4.7% at a
maximum.
Then, we have investigated the behavior and performance of the fault-tolerant
mSPMD programming execution when errors occur. Instead of waiting for real
errors, we have inserted fake errors that stop heartbeat messages from remote
programs randomly with a certain error probability computed by an expected
0
100
200
300
400
500
600
700
64
256
512
1024
Execution time (sec)
(procs/task)
20480x20480 Matrix, 1024 processes in total
01x01 woft
02x02 woft
04x04 woft
08x08 woft
01x01 w/ft
02x02 w/ft
04x04 w/ft
08x08 w/ft
Fig. 15 Execution time with and without FT for the number of cores for each task. The graph
legends show the number of blocks
235
# of blocks 1 × 1
2 × 2
4 × 4 8 × 8
Block size 20,480 2 10,240 2 5120 2 2560 2
1024 cores (64 nodes) are used for each workflow, and 64–1024 cores are
assigned for each task in a workflow.
Firstly, we have considered the overhead of the heartbeat messages used to
detect errors in remote programs. Figure 15 shows the performance of the normal
and fault-tolerant mSPMD programming executions using between 64 and 1024
compute cores per task. The dotted lines are the results of fault-tolerant mSPMD
programming executions, and the solid lines are those of the regular mSPMD
programming executions. As shown in the figure, the best combination of the
number of blocks and the number of processes per task is 4 × 4 blocks and 512
processes for both cases of with and without fault-tolerance support. The overhead
of using a heartbeat message is very small and is 2.3% on average and 4.7% at a
maximum.
Then, we have investigated the behavior and performance of the fault-tolerant
mSPMD programming execution when errors occur. Instead of waiting for real
errors, we have inserted fake errors that stop heartbeat messages from remote
programs randomly with a certain error probability computed by an expected
0
100
200
300
400
500
600
700
64
256
512
1024
Execution time (sec)
(procs/task)
20480x20480 Matrix, 1024 processes in total
01x01 woft
02x02 woft
04x04 woft
08x08 woft
01x01 w/ft
02x02 w/ft
04x04 w/ft
08x08 w/ft
Fig. 15 Execution time with and without FT for the number of cores for each task. The graph
legends show the number of blocks
