Theor Chem Acc (2015) 134:102
1 3
several thousands of GPU threads. The role of the computational kernel is to perform operations summarized by
Eqs. ( 8 )–( 16 ). Once the sub-matrix B * is generated in the
GPU memory, we carry out the matrix multiplication ( 25 )
by the use of the CUBLAS library. Results are then transferred into the CPU memory space (RAM of the computer
that hosts the GPU card) and saved on the disk. Size of the
segment computed in one pass is determined by the size of
the GPU memory, which was 6 GB in the present case. As
can be seen, the proposed algorithm for the derivatives of
the exchange integrals requires, in general, very large data
blocks denoted as B * in Eq. ( 17 ). However, in the presented
GPU implementation these data blocks are generated on
the GPU card, they are used on the GPU device, and they
are deleted from the GPU memory once they are of no use
for the following calculations. Therefore, there is no overhead time connected with their manipulation.
From our personal experience, an exploitation of GPU
for its effi cient use is not an easy and straightforward task.
Large parts of the programs designed for CPU must be
rewritten, and algorithms must be rethought. For example, an alignment of data in the GPU memory plays a very
important role in the speed of the program. A GPU implementation often requires (as in the present case) programming of a GPU kernel that is designed to be launched in
thousands of instances. The available programing and especially the debugging tools are still very limited. Even, with
all these diffi culties, the GPGPU implementations presently
undergo very rapid development in software and hardware
and thus they offer a promising environment for computational quantum chemistry and physics.
4 Performance and accuracy
For testing purposes, we performed calculations for all vibrational modes of cyclopropane, benzene and adamantane
in two different ways. In the fi rst approach, the exchange
integrals and their derivatives were evaluated rigorously by
means of complex Shavitt functions F n ( z ) by the original
(CPU parallelized) version of the program [ 1 , 3 ]. We used 12
CPU cores, because on a single core the calculations would
be excessively long, in particular for adamantane.
In the second approach, a single CPU was used but
with the help of GPU for factorization given by Eqs. ( 3 )
and ( 17 ). The entries in Table 1 are the respective timings
for the evaluation of all integrals and their derivatives that
were needed for computation of differential cross sections
for all vibrational modes. The CPU execution times, t 0 ,
shown in Table 1 on the fi rst line were obtained by using
12 cores of Intel Xeon 3 GHz workstation. The entries on
the second line are execution times on a single CPU core
code where the evaluation steps shown in Eqs. ( 3 ) and ( 17 )
were offl oaded to Tesla card K20m with 2496 cuda cores
clocked at 710 MHz. Single-core rigorous calculations for
benzene and adamantane would be excessively long, and
therefore, they were only performed with our parallelized
version of the program for 12 CPU’s. Speedups are therefore expressed as ratios t 0 / t FT . They may be multiplied by
a factor of 11.1, which was the ratio of timings for rigorous calculations we obtained for cyclopropane on 1 and 12
CPU’s.
For checking the loss in accuracy connected with the
numerical discretization of the 1/ r term, we selected the
cyclopropane molecule and derivatives with respect to the
C 2 y atomic coordinate, as it is defi ned in Fig. 1 . Testing
was done for many combinations of k 1 and k 2 , keeping the
k 2 vectors fi xed at 12 orientations, whereas the k 1 vectors
were varied.
The result of the testing is shown graphically in Figs. 2
and 3 . Figure 2 shows a case when variation of k 1 is limited
to small values of | k 1 |. For plots, we selected the regions
of k 1 and k 2 combinations, where the error was the largest.
With a smaller numerical quadrature used in Eq. ( 2 ), the
maximum error was found for k 2 ≡ (0; 10; 0) and it was in
the range of 10
−4 a.u. The purpose of this fi gure is to show
that on increasing the number of grid points from 3402 to
31,516, the maximum error was reduced to the range of
microhartrees.
Figure 3 shows the only region where the factorization
method does not work well. If both k 1 and k 2 are large in
absolute value and parallel or close to be parallel in orientation, the error is in the millihartree range. On increasing the
number of grid points from 3402 to 31,516, the maximum
error was reduced, but only to the range of 10
−4 a.u. If a
higher accuracy is requested, this small number of integrals
and their derivatives must be evaluated by some other more
rigorous method. It is unlikely that the mere use of the
double-precision arithmetics would solve the problem. We
reached the limit in accuracy of single-precision arithmetics that is suffi cient for our scattering codes, and therefore,
we did not attempt to optimize the numerical FT quadrature
further.
Table 1 CPU time (in seconds) and speedup in calculations of a
preselected set of exchange integrals ( k 1 | V ex | k 2 ) and their derivatives
∂/∂ A x ( k 1 | V ex | k 2 ) (denoted as ∂/∂y in Sect. 2 ), needed for computation of
vibrational cross sections of cyclopropane, benzene and adamantane
a Calculation by means of complex Shavitt functions F n ( z )
b Calculation by means of the Fourier transform of 1/ r
C 3 H 6
C 6 H 6
C 10 H 16
Rigorous calculation
a , 12 CPU’s, t 0
19,894 67,330 515,100
FT calculation
b , GPU used, 1 CPU, t FT
1398
3610
19,460
Speedup, defi ned as ( t 0 / t FT )
14
19
26
19
Reprinted from the journal
1 3
several thousands of GPU threads. The role of the computational kernel is to perform operations summarized by
Eqs. ( 8 )–( 16 ). Once the sub-matrix B * is generated in the
GPU memory, we carry out the matrix multiplication ( 25 )
by the use of the CUBLAS library. Results are then transferred into the CPU memory space (RAM of the computer
that hosts the GPU card) and saved on the disk. Size of the
segment computed in one pass is determined by the size of
the GPU memory, which was 6 GB in the present case. As
can be seen, the proposed algorithm for the derivatives of
the exchange integrals requires, in general, very large data
blocks denoted as B * in Eq. ( 17 ). However, in the presented
GPU implementation these data blocks are generated on
the GPU card, they are used on the GPU device, and they
are deleted from the GPU memory once they are of no use
for the following calculations. Therefore, there is no overhead time connected with their manipulation.
From our personal experience, an exploitation of GPU
for its effi cient use is not an easy and straightforward task.
Large parts of the programs designed for CPU must be
rewritten, and algorithms must be rethought. For example, an alignment of data in the GPU memory plays a very
important role in the speed of the program. A GPU implementation often requires (as in the present case) programming of a GPU kernel that is designed to be launched in
thousands of instances. The available programing and especially the debugging tools are still very limited. Even, with
all these diffi culties, the GPGPU implementations presently
undergo very rapid development in software and hardware
and thus they offer a promising environment for computational quantum chemistry and physics.
4 Performance and accuracy
For testing purposes, we performed calculations for all vibrational modes of cyclopropane, benzene and adamantane
in two different ways. In the fi rst approach, the exchange
integrals and their derivatives were evaluated rigorously by
means of complex Shavitt functions F n ( z ) by the original
(CPU parallelized) version of the program [ 1 , 3 ]. We used 12
CPU cores, because on a single core the calculations would
be excessively long, in particular for adamantane.
In the second approach, a single CPU was used but
with the help of GPU for factorization given by Eqs. ( 3 )
and ( 17 ). The entries in Table 1 are the respective timings
for the evaluation of all integrals and their derivatives that
were needed for computation of differential cross sections
for all vibrational modes. The CPU execution times, t 0 ,
shown in Table 1 on the fi rst line were obtained by using
12 cores of Intel Xeon 3 GHz workstation. The entries on
the second line are execution times on a single CPU core
code where the evaluation steps shown in Eqs. ( 3 ) and ( 17 )
were offl oaded to Tesla card K20m with 2496 cuda cores
clocked at 710 MHz. Single-core rigorous calculations for
benzene and adamantane would be excessively long, and
therefore, they were only performed with our parallelized
version of the program for 12 CPU’s. Speedups are therefore expressed as ratios t 0 / t FT . They may be multiplied by
a factor of 11.1, which was the ratio of timings for rigorous calculations we obtained for cyclopropane on 1 and 12
CPU’s.
For checking the loss in accuracy connected with the
numerical discretization of the 1/ r term, we selected the
cyclopropane molecule and derivatives with respect to the
C 2 y atomic coordinate, as it is defi ned in Fig. 1 . Testing
was done for many combinations of k 1 and k 2 , keeping the
k 2 vectors fi xed at 12 orientations, whereas the k 1 vectors
were varied.
The result of the testing is shown graphically in Figs. 2
and 3 . Figure 2 shows a case when variation of k 1 is limited
to small values of | k 1 |. For plots, we selected the regions
of k 1 and k 2 combinations, where the error was the largest.
With a smaller numerical quadrature used in Eq. ( 2 ), the
maximum error was found for k 2 ≡ (0; 10; 0) and it was in
the range of 10
−4 a.u. The purpose of this fi gure is to show
that on increasing the number of grid points from 3402 to
31,516, the maximum error was reduced to the range of
microhartrees.
Figure 3 shows the only region where the factorization
method does not work well. If both k 1 and k 2 are large in
absolute value and parallel or close to be parallel in orientation, the error is in the millihartree range. On increasing the
number of grid points from 3402 to 31,516, the maximum
error was reduced, but only to the range of 10
−4 a.u. If a
higher accuracy is requested, this small number of integrals
and their derivatives must be evaluated by some other more
rigorous method. It is unlikely that the mere use of the
double-precision arithmetics would solve the problem. We
reached the limit in accuracy of single-precision arithmetics that is suffi cient for our scattering codes, and therefore,
we did not attempt to optimize the numerical FT quadrature
further.
Table 1 CPU time (in seconds) and speedup in calculations of a
preselected set of exchange integrals ( k 1 | V ex | k 2 ) and their derivatives
∂/∂ A x ( k 1 | V ex | k 2 ) (denoted as ∂/∂y in Sect. 2 ), needed for computation of
vibrational cross sections of cyclopropane, benzene and adamantane
a Calculation by means of complex Shavitt functions F n ( z )
b Calculation by means of the Fourier transform of 1/ r
C 3 H 6
C 6 H 6
C 10 H 16
Rigorous calculation
a , 12 CPU’s, t 0
19,894 67,330 515,100
FT calculation
b , GPU used, 1 CPU, t FT
1398
3610
19,460
Speedup, defi ned as ( t 0 / t FT )
14
19
26
19
Reprinted from the journal
