XcalableMP 2.0 and Future Directions
257
for each operation performed on all threads, because tasks executed on threads
differ each time the program is executed. The “wait” in the breakdown represents
the waiting time of the thread, including the global synchronization. The “comm”
indicates the time from the start of the communication to the end. In Fig. 8a, the
“Task” version shows a better performance than the barrier-based “Parallel Loop”
implementation. The reason that the “Task” version outperforms the “Parallel Loop”
version is that the global synchronization uses a higher cost for the work sharing of
loops and among tasks, as shown in Fig. 8b. The relative performance of the “Task”
version compared with the “Parallel Loop” version is 123% (Fig. 8).
3.5 Communication Optimization for Manycore Clusters
In the global task parallel programming model, the communication may happen
at each pair of tasks between nodes. In order to enable the communication
in multithreaded environment, we may use MPI_THREAD_MULTIPLE as the
MPI thread-safety level, because tasks executed on threads may communicate
simultaneously. We have examined the basic performance of multithreaded communications by using the Ping-Pong benchmark. This benchmark is based on the
OSU Micro-Benchmarks 5.3.2 [12] developed by the Ohio State University. we
also show the aggregated bandwidth when multiple threads (i.e., two, four, or eight
threads) communicate at the same time. Figure 9 illustrates the communication
performance on the Oakforest-PACS system. The performance of multithreaded
communication with MPI_THREAD_MULTIPLE degrades compared to a singlethreaded communication as the number of threads increases. As with the result
on the Oakforest-PACS system, the performance of communication on a single thread is better compared to that for multithreaded communication with
MPI_THREAD_MULTIPLE. Therefore, the communication performance may be
improved if all communications are delegated to the communication thread. To
delegate the communications to a single thread, we create a global queue that is
accessible by all threads, so that the tasks enqueue the communication requests into
this queue and wait for the communication to complete. Meanwhile, the communication thread dequeues the requests for communication to perform the requested
communications, and checks the communication completion. The communication
thread executes only the communication, and the other threads perform computation
tasks.
Figure 8 shows the performance and breakdown of blocked Cholesky factorization with the communication optimization denoted as “Task (opt).” The “Task
(opt)” version of blocked Cholesky factorization performs better than the multitasking execution with MPI_THREAD_MULTIPLE. The reason for this is that
the communication time decreases compared with the “Task” version, as shown in
Figs. 8, because of the use of the communication thread. The relative performances
compared with the barrier-based “Parallel Loop” implementation improve to 138%
on the Oakforest-PACS systems.
257
for each operation performed on all threads, because tasks executed on threads
differ each time the program is executed. The “wait” in the breakdown represents
the waiting time of the thread, including the global synchronization. The “comm”
indicates the time from the start of the communication to the end. In Fig. 8a, the
“Task” version shows a better performance than the barrier-based “Parallel Loop”
implementation. The reason that the “Task” version outperforms the “Parallel Loop”
version is that the global synchronization uses a higher cost for the work sharing of
loops and among tasks, as shown in Fig. 8b. The relative performance of the “Task”
version compared with the “Parallel Loop” version is 123% (Fig. 8).
3.5 Communication Optimization for Manycore Clusters
In the global task parallel programming model, the communication may happen
at each pair of tasks between nodes. In order to enable the communication
in multithreaded environment, we may use MPI_THREAD_MULTIPLE as the
MPI thread-safety level, because tasks executed on threads may communicate
simultaneously. We have examined the basic performance of multithreaded communications by using the Ping-Pong benchmark. This benchmark is based on the
OSU Micro-Benchmarks 5.3.2 [12] developed by the Ohio State University. we
also show the aggregated bandwidth when multiple threads (i.e., two, four, or eight
threads) communicate at the same time. Figure 9 illustrates the communication
performance on the Oakforest-PACS system. The performance of multithreaded
communication with MPI_THREAD_MULTIPLE degrades compared to a singlethreaded communication as the number of threads increases. As with the result
on the Oakforest-PACS system, the performance of communication on a single thread is better compared to that for multithreaded communication with
MPI_THREAD_MULTIPLE. Therefore, the communication performance may be
improved if all communications are delegated to the communication thread. To
delegate the communications to a single thread, we create a global queue that is
accessible by all threads, so that the tasks enqueue the communication requests into
this queue and wait for the communication to complete. Meanwhile, the communication thread dequeues the requests for communication to perform the requested
communications, and checks the communication completion. The communication
thread executes only the communication, and the other threads perform computation
tasks.
Figure 8 shows the performance and breakdown of blocked Cholesky factorization with the communication optimization denoted as “Task (opt).” The “Task
(opt)” version of blocked Cholesky factorization performs better than the multitasking execution with MPI_THREAD_MULTIPLE. The reason for this is that
the communication time decreases compared with the “Task” version, as shown in
Figs. 8, because of the use of the communication thread. The relative performances
compared with the barrier-based “Parallel Loop” implementation improve to 138%
on the Oakforest-PACS systems.
