5.8 Memory Management and Two-Dimensional Partitioning
223
which removes all the angular momenta J = J . The J quantum number
conservation has been checked to be numerically precise up to at least 10 −10 in
Gamow shell model applications.
The standard formulation ˆ
J 2 = ˆ
J − ˆ
J + + M(M + 1) is utilized to apply ˆ
J 2 ,
where ˆ
J ± is a sparse one-body operator. The computation of ˆ
J 2 times vector is thus
fast. The main problem, however, is that the direct application of 2D partitioning is
rather inefficient, that due to the block structure of the ˆ
J ± operator matrix. Indeed,
as ˆ
J ± cannot connect two different configurations, it is entirely contained in a
diagonal square of its matrix, except for the very few Slater determinants situated in
two adjacent nodes and belonging to the same configurations. Therefore, a naive
application of 2D partitioning to the ˆ
J ± matrix would have the diagonal nodes
associated to ˆ
J ± to contain virtually all the ˆ
J ± matrix, so that all the other nodes
would be inactive. Hence, the ˆ
J ± routines have been parallelized with a hybrid
1D/2D method, where a column of the ˆ
J ± matrix is stored on one node. Gamow
shell model vectors are then divided into N parts, similarly to the 2D partitioning
scheme, so that the output Gamow shell model vector can still be reduced on each
node after application of the ˆ
J ± operator. Considering that each node contains a
part of the diagonal squares forming the ˆ
J ± matrix, the memory distribution of the
ˆ
J ± matrix is well balanced.
ˆ
J ± is not symmetric because it connects Slater determinants of total angular
momentum projection M and M ± 1. Hence, one does not have to account for
symmetry in the parallelization of ˆ
J ± times vector, contrary to the 2D partitioning
of the ˆ
H matrix (see Sect. 5.8.3).
In fact, the fundamental problem here is message passing interface data transfer.
Indeed, the number of angular momenta J to suppress (see Eq. (5.71)) is of the
order of 50, so that, naively, one would have to do two message passing interface
transfers of a full Gamow shell model vector per ˆ
J ± application. This would lead
to 200 message passing interface transfers of a full Gamow shell model vector per
ˆ
P J application, which is prohibitively large. In order to solve this problem, one
takes advantage of the block structure of the ˆ
J ± matrix. Indeed, the necessary
message passing interface transfers involve only neighboring nodes in Gamow shell
model vectors. Thus, their total volume is that of the nonzero components of a
configuration in between two nodes, and not the d dimension, so that the amount
of data to transfer is tractable. Moreover, the application of ˆ
P J occurs only a few
times, because the J quantum number is most often numerically conserved after a ˆ
H
application. Therefore, the projection on J quantum number does not significantly
slow down the implementation of eigenvectors.
In order to know whether or not it is necessary to apply ˆ
P J , one has to test if
a Gamow shell model vector is an eigenstate of the ˆ
J 2 operator. If M = J , the
action of ˆ
J + on a J -coupled Gamow shell model vector provides zero, because one
cannot generate a J -coupled many-body state whose angular momentum projection
is J + 1. Hence, ˆ
P J has to be applied if the action of ˆ
J + on a Gamow shell model
vector provides with a nonzero vector. This operation does not create any sizable
computational burden because the ˆ
J + matrix is very sparse, so that the time taken by
223
which removes all the angular momenta J = J . The J quantum number
conservation has been checked to be numerically precise up to at least 10 −10 in
Gamow shell model applications.
The standard formulation ˆ
J 2 = ˆ
J − ˆ
J + + M(M + 1) is utilized to apply ˆ
J 2 ,
where ˆ
J ± is a sparse one-body operator. The computation of ˆ
J 2 times vector is thus
fast. The main problem, however, is that the direct application of 2D partitioning is
rather inefficient, that due to the block structure of the ˆ
J ± operator matrix. Indeed,
as ˆ
J ± cannot connect two different configurations, it is entirely contained in a
diagonal square of its matrix, except for the very few Slater determinants situated in
two adjacent nodes and belonging to the same configurations. Therefore, a naive
application of 2D partitioning to the ˆ
J ± matrix would have the diagonal nodes
associated to ˆ
J ± to contain virtually all the ˆ
J ± matrix, so that all the other nodes
would be inactive. Hence, the ˆ
J ± routines have been parallelized with a hybrid
1D/2D method, where a column of the ˆ
J ± matrix is stored on one node. Gamow
shell model vectors are then divided into N parts, similarly to the 2D partitioning
scheme, so that the output Gamow shell model vector can still be reduced on each
node after application of the ˆ
J ± operator. Considering that each node contains a
part of the diagonal squares forming the ˆ
J ± matrix, the memory distribution of the
ˆ
J ± matrix is well balanced.
ˆ
J ± is not symmetric because it connects Slater determinants of total angular
momentum projection M and M ± 1. Hence, one does not have to account for
symmetry in the parallelization of ˆ
J ± times vector, contrary to the 2D partitioning
of the ˆ
H matrix (see Sect. 5.8.3).
In fact, the fundamental problem here is message passing interface data transfer.
Indeed, the number of angular momenta J to suppress (see Eq. (5.71)) is of the
order of 50, so that, naively, one would have to do two message passing interface
transfers of a full Gamow shell model vector per ˆ
J ± application. This would lead
to 200 message passing interface transfers of a full Gamow shell model vector per
ˆ
P J application, which is prohibitively large. In order to solve this problem, one
takes advantage of the block structure of the ˆ
J ± matrix. Indeed, the necessary
message passing interface transfers involve only neighboring nodes in Gamow shell
model vectors. Thus, their total volume is that of the nonzero components of a
configuration in between two nodes, and not the d dimension, so that the amount
of data to transfer is tractable. Moreover, the application of ˆ
P J occurs only a few
times, because the J quantum number is most often numerically conserved after a ˆ
H
application. Therefore, the projection on J quantum number does not significantly
slow down the implementation of eigenvectors.
In order to know whether or not it is necessary to apply ˆ
P J , one has to test if
a Gamow shell model vector is an eigenstate of the ˆ
J 2 operator. If M = J , the
action of ˆ
J + on a J -coupled Gamow shell model vector provides zero, because one
cannot generate a J -coupled many-body state whose angular momentum projection
is J + 1. Hence, ˆ
P J has to be applied if the action of ˆ
J + on a Gamow shell model
vector provides with a nonzero vector. This operation does not create any sizable
computational burden because the ˆ
J + matrix is very sparse, so that the time taken by
