Theor Chem Acc (2015) 134:102
1 3
manageable approach to polyatomic molecules. Hence, the
U 10 matrix elements for the i th normal mode can be taken as
the derivative of the electron–molecule interaction potential
with respect to the i th normal coordinate,
The calculations were performed for the incident electron
energy of 10 eV. The one-electron A and B terms were evaluated for 62 680 806 k 1 , k 2 pairs, with the absolute value
of k vectors within the range from 0 to 17 a.u. For more
details, we refer to the quoted papers.
3.2 Hartree–Fock calculations
For Hartree–Fock calculations of cyclopropane, benzene
and adamantane, we used the Gaussian valence-shell double-zeta-and-polarization (9 s 5 p /4 s 1 p )/[3 s 2 p 1 d /2 s 1 p ] basis
set of Dunning and Hay [ 8 ]. The geometries of the three
molecules were optimized with this DZP basis set and
then used for analytical evaluation of normal modes, harmonic frequencies, dipole moment and its derivatives, and
density matrix and its derivatives with respect to atomic
coordinates.
3.3 Numerical quadrature
In the x
y notation of Chien and Gill [ 9 ], the quadrature
used for UGT terms in Eq. ( 18 ) may be denoted as 6
12 12
1
38
1 590
2 146
1 194
2 302
1 434
2 770
7 2030
1 , indicating that
x point angular Lebedev grid [ 10 , 11 ] was used for y successive radial points of the Gauss–Legendre quadrature.
The total number of radial points was 31. The 31st radial
point with 770 angular grid points was assigned to k 0 of
the incoming electron. The total number of grid points
was 11,196. The quadrature used for factorization of the
1/ r operator in Eq. ( 2 ) was of the Laguerre–Lebedev type.
After some experimentation, we found that a scaled adaptive Laguerre quadrature with 15 grid points can be used
as an universal radial quadrature for low-energy electron–
molecule scattering calculations. The radial points were
obtained as scaled Laguerre grid points
where x 15, i is the i th grid point in the 15-point Laguerre
quadrature and the scaling factor R was obtained as
with the fi xed R max = 11 a.u. The weights were also scaled
as
where w i ’s are weights of Laguerre grid points. The scaled
radial points k i were ordered in fi ve ranges (0, 0.2), (0.2,
(21)
U 10 = 1/
√
2 ∂U/∂q i
(22)
k i = x 15,i R,
(23)
R = R max /x 15,15 ,
(24)
ω i = w i e
x i R,
0.4), (0.4, 0.8), (0.8, 1.0) and (1.0, 11.0) and assigned to
Lebedev angular quadratures with 26, 50, 86, 194 and 302
grid points, respectively. The total number of grid points
was 3402. In the x
y notation of Chien and Gill, the quadrature may be denoted as 26
2 50
1 86
1 194
1 302
10 . Some calculations were performed with a larger quadrature with a
total number of 31,516 grid points, to see if the accuracy of
the calculated integrals and derivatives can be improved to
the μHartree range. R max was fi xed at 20, and 50 Laguerre
radial points were ordered in six ranges (0, 0.5), (0.5,
3.0), (3.0, 7.5), (7.5, 9.5), (9.5, 15.0) and (15.0, 20.0) as
50
9 302
14 590
14 770
1 2030
8 302
4 . It should be noted that
this quadrature with 31,516 grid points is used here only
for testing. For practical calculations, it is too large as it
increases the computational time so much that the merit of
time-saving is lost.
3.4 GPU implementation of the method
The bottleneck of the computational scheme, as it is represented by Eqs. ( 8 )–( 17 ), is the calculation and storage
of the B terms defi ned by Eq. ( 16 ). For a somewhat larger
molecular system, the size of the B array may quickly
exceed common disk capacities; moreover, a manipulation with such a large data object (represented by necessary
input/output operations) quickly becomes a dominant hindrance in the calculations. For example, in case of adamantane, which serves as one of the benchmark molecules in
the present paper, the physical size of the B array that is
necessary to form the U 10 matrix in Eqs. ( 18 ), ( 19 ), is larger
than 1 TB of data. Therefore, we do not recommend a typical CPU implementation of presented method. Instead, we
propose an implementation in which the elements of the
B array are not saved on any storage device. In fact, they
never even enter the memory space of the computer. This
idea is implemented as follows.
Let M be a collective index for plane-wave k 1 and the
nuclear coordinate we differentiate with respect to. Thus,
M represents two upper indices of the B array in Eq. ( 16 ).
Let N be a contraction index in Eq. ( 17 ); i.e., N represents
a collective index for occupied orbital i and plane-wave
q ( 4 ), denoted as lower indices of B array in Eq. ( 16 ).
Then, the fi rst term in Eq. ( 17 ) represents a simple matrix
multiplication
while the second term is a Hermitian conjugate of the
fi rst term. Since the array B is too large to be stored in
the GPU memory, the multiplication ( 25 ) must be done in
more passes. For each pass, we fi rst calculate a sub-matrix
B
* for a subset of rows M 1 – M 2 . This computation is carried out by a simultaneous launch of a CUDA kernel on
(25)
N
B
∗
(M, N)A
N, M
,
18
Reprinted from the journal
1 3
manageable approach to polyatomic molecules. Hence, the
U 10 matrix elements for the i th normal mode can be taken as
the derivative of the electron–molecule interaction potential
with respect to the i th normal coordinate,
The calculations were performed for the incident electron
energy of 10 eV. The one-electron A and B terms were evaluated for 62 680 806 k 1 , k 2 pairs, with the absolute value
of k vectors within the range from 0 to 17 a.u. For more
details, we refer to the quoted papers.
3.2 Hartree–Fock calculations
For Hartree–Fock calculations of cyclopropane, benzene
and adamantane, we used the Gaussian valence-shell double-zeta-and-polarization (9 s 5 p /4 s 1 p )/[3 s 2 p 1 d /2 s 1 p ] basis
set of Dunning and Hay [ 8 ]. The geometries of the three
molecules were optimized with this DZP basis set and
then used for analytical evaluation of normal modes, harmonic frequencies, dipole moment and its derivatives, and
density matrix and its derivatives with respect to atomic
coordinates.
3.3 Numerical quadrature
In the x
y notation of Chien and Gill [ 9 ], the quadrature
used for UGT terms in Eq. ( 18 ) may be denoted as 6
12 12
1
38
1 590
2 146
1 194
2 302
1 434
2 770
7 2030
1 , indicating that
x point angular Lebedev grid [ 10 , 11 ] was used for y successive radial points of the Gauss–Legendre quadrature.
The total number of radial points was 31. The 31st radial
point with 770 angular grid points was assigned to k 0 of
the incoming electron. The total number of grid points
was 11,196. The quadrature used for factorization of the
1/ r operator in Eq. ( 2 ) was of the Laguerre–Lebedev type.
After some experimentation, we found that a scaled adaptive Laguerre quadrature with 15 grid points can be used
as an universal radial quadrature for low-energy electron–
molecule scattering calculations. The radial points were
obtained as scaled Laguerre grid points
where x 15, i is the i th grid point in the 15-point Laguerre
quadrature and the scaling factor R was obtained as
with the fi xed R max = 11 a.u. The weights were also scaled
as
where w i ’s are weights of Laguerre grid points. The scaled
radial points k i were ordered in fi ve ranges (0, 0.2), (0.2,
(21)
U 10 = 1/
√
2 ∂U/∂q i
(22)
k i = x 15,i R,
(23)
R = R max /x 15,15 ,
(24)
ω i = w i e
x i R,
0.4), (0.4, 0.8), (0.8, 1.0) and (1.0, 11.0) and assigned to
Lebedev angular quadratures with 26, 50, 86, 194 and 302
grid points, respectively. The total number of grid points
was 3402. In the x
y notation of Chien and Gill, the quadrature may be denoted as 26
2 50
1 86
1 194
1 302
10 . Some calculations were performed with a larger quadrature with a
total number of 31,516 grid points, to see if the accuracy of
the calculated integrals and derivatives can be improved to
the μHartree range. R max was fi xed at 20, and 50 Laguerre
radial points were ordered in six ranges (0, 0.5), (0.5,
3.0), (3.0, 7.5), (7.5, 9.5), (9.5, 15.0) and (15.0, 20.0) as
50
9 302
14 590
14 770
1 2030
8 302
4 . It should be noted that
this quadrature with 31,516 grid points is used here only
for testing. For practical calculations, it is too large as it
increases the computational time so much that the merit of
time-saving is lost.
3.4 GPU implementation of the method
The bottleneck of the computational scheme, as it is represented by Eqs. ( 8 )–( 17 ), is the calculation and storage
of the B terms defi ned by Eq. ( 16 ). For a somewhat larger
molecular system, the size of the B array may quickly
exceed common disk capacities; moreover, a manipulation with such a large data object (represented by necessary
input/output operations) quickly becomes a dominant hindrance in the calculations. For example, in case of adamantane, which serves as one of the benchmark molecules in
the present paper, the physical size of the B array that is
necessary to form the U 10 matrix in Eqs. ( 18 ), ( 19 ), is larger
than 1 TB of data. Therefore, we do not recommend a typical CPU implementation of presented method. Instead, we
propose an implementation in which the elements of the
B array are not saved on any storage device. In fact, they
never even enter the memory space of the computer. This
idea is implemented as follows.
Let M be a collective index for plane-wave k 1 and the
nuclear coordinate we differentiate with respect to. Thus,
M represents two upper indices of the B array in Eq. ( 16 ).
Let N be a contraction index in Eq. ( 17 ); i.e., N represents
a collective index for occupied orbital i and plane-wave
q ( 4 ), denoted as lower indices of B array in Eq. ( 16 ).
Then, the fi rst term in Eq. ( 17 ) represents a simple matrix
multiplication
while the second term is a Hermitian conjugate of the
fi rst term. Since the array B is too large to be stored in
the GPU memory, the multiplication ( 25 ) must be done in
more passes. For each pass, we fi rst calculate a sub-matrix
B
* for a subset of rows M 1 – M 2 . This computation is carried out by a simultaneous launch of a CUDA kernel on
(25)
N
B
∗
(M, N)A
N, M
,
18
Reprinted from the journal
