84
M. Nakao and H. Murai
5 Performance Evaluation
In order to evaluate the performance of XMP, we implemented the HPC Challenge (HPCC) benchmark (https://icl.utk.edu/hpcc/), namely, EP STREAM Triad
(STREAM), High-Performance Linpack (HPL), Global fast Fourier transform
(FFT), and RandomAccess [5]. While the HPCC benchmark is used to evaluate
multiple attributes of HPC systems, the benchmark is also useful to evaluate the
properties of a parallel language. The HPCC benchmark was used at the HPCC
Award Competition (https://www.hpcchallenge.org). The HPCC Award Competition consists of two classes. While the purpose of class 1 is to evaluate the
performance of a machine, the purpose of class 2 is to evaluate both the productivity
and performance of a parallel programming language. XMP won the class 2 prizes
in 2013 and 2014.
5.1 Experimental Environment
For performance evaluation, this section uses 16,384 compute nodes on the K computer and 128 compute nodes on a Cray CS300 system named “the COMA system.”
Tables 1 and 2 show the hardware specifications and software environments.
For comparison purposes, this section also evaluates the HPCC benchmark in C
language and MPI library. We execute STREAM, HPL, and FFT with eight threads
per process on each CPU of the K computer, and with ten threads per process on
each CPU of the COMA system. Since RandomAccess is not parallelized with
threads and can be executed by the power of only two processes, we execute it
with eight processes on each CPU of both systems.
The specification of HPCC Award Competition class 2 defines the minimum
problem size for each benchmark. While the main array of HPL should occupy
at least half of the system memory, the main arrays of STREAM, FFT, and
RandomAccess should occupy at least a quarter of the system memory. We set each
Table 1 Experimental environment for the K computer
CPU
SPARC64 VIIIfx 2.0 GHz, 8 Cores
Memory
DDR3 SDRAM 16 GB, 64 GB/s
Network
Torus fusion six-dimensional mesh/torus network, 5 GB/s × 10
Library
Fujitsu Compiler K-1.2.0-19, Fujitsu MPI K-1.2.0-19, Fujitsu SSLII K-1.2.0-19
Table 2 Experimental environment for the COMA system
CPU
Xeon E5-2670v2, 2.5 GHz (Turbo Boost 3.3 GHz), 10 Cores × 2CPUs
Memory
DDR3 SDRAM 64 GB, 119.4 GB/s (= 59.7 GB/s × 2 CPUs)
Network
InfiniBand FDR, fat-tree, 7 GB/s
Library
Intel Compiler 15.0.5, Intel MPI 5.1.1, GASNet 1.26.0, Intel MKL 11.2.4
M. Nakao and H. Murai
5 Performance Evaluation
In order to evaluate the performance of XMP, we implemented the HPC Challenge (HPCC) benchmark (https://icl.utk.edu/hpcc/), namely, EP STREAM Triad
(STREAM), High-Performance Linpack (HPL), Global fast Fourier transform
(FFT), and RandomAccess [5]. While the HPCC benchmark is used to evaluate
multiple attributes of HPC systems, the benchmark is also useful to evaluate the
properties of a parallel language. The HPCC benchmark was used at the HPCC
Award Competition (https://www.hpcchallenge.org). The HPCC Award Competition consists of two classes. While the purpose of class 1 is to evaluate the
performance of a machine, the purpose of class 2 is to evaluate both the productivity
and performance of a parallel programming language. XMP won the class 2 prizes
in 2013 and 2014.
5.1 Experimental Environment
For performance evaluation, this section uses 16,384 compute nodes on the K computer and 128 compute nodes on a Cray CS300 system named “the COMA system.”
Tables 1 and 2 show the hardware specifications and software environments.
For comparison purposes, this section also evaluates the HPCC benchmark in C
language and MPI library. We execute STREAM, HPL, and FFT with eight threads
per process on each CPU of the K computer, and with ten threads per process on
each CPU of the COMA system. Since RandomAccess is not parallelized with
threads and can be executed by the power of only two processes, we execute it
with eight processes on each CPU of both systems.
The specification of HPCC Award Competition class 2 defines the minimum
problem size for each benchmark. While the main array of HPL should occupy
at least half of the system memory, the main arrays of STREAM, FFT, and
RandomAccess should occupy at least a quarter of the system memory. We set each
Table 1 Experimental environment for the K computer
CPU
SPARC64 VIIIfx 2.0 GHz, 8 Cores
Memory
DDR3 SDRAM 16 GB, 64 GB/s
Network
Torus fusion six-dimensional mesh/torus network, 5 GB/s × 10
Library
Fujitsu Compiler K-1.2.0-19, Fujitsu MPI K-1.2.0-19, Fujitsu SSLII K-1.2.0-19
Table 2 Experimental environment for the COMA system
CPU
Xeon E5-2670v2, 2.5 GHz (Turbo Boost 3.3 GHz), 10 Cores × 2CPUs
Memory
DDR3 SDRAM 64 GB, 119.4 GB/s (= 59.7 GB/s × 2 CPUs)
Network
InfiniBand FDR, fat-tree, 7 GB/s
Library
Intel Compiler 15.0.5, Intel MPI 5.1.1, GASNet 1.26.0, Intel MKL 11.2.4
