In order to improve the estimates of the protein motion correlations a statistical analysis over multiple MD simulations should be
performed. The correlation coefficients can be computed for each
independent MD trajectory (or as a single computation of concatenated independent trajectories) and then averaged out. As mentioned above, the correlation coefficients (including r[x i ,x j ]) can be
computed using the GROMACS software, which requires the
topology file (containing all the necessary information to compute
the forces and velocities for the MD) to be in the GROMACS
format. This topology file format can be obtained from other
topology formats (e.g., from AMBER topology file) using free
software, such as ParmEd [46].
3 Methods
3.1 Dynamical
Weighted Network
The pair-correlations coefficients defined above measure how correlated are the motion of two residues recorded during a MD
simulation. When employing the mutual information measure,
i.e., the generalized r MI [x i ,x j ] coefficients, the pair-correlations
are directly related to the exchange of information between two
residues and can then be used to weight a graph that models the
information exchange within a protein according to its dynamics,
i.e., to build a dynamical weighted network. First, a protein graph
needs to be defined. Since the correlation coefficients are computed
using the C α atomic positions fluctuations, each node of the graph
is associated to each C α of the protein representing each amino acid
residue in the primary sequence, see Fig. 2a. To complete the graph,
one needs to build an adjacency matrix, A, i.e., a matrix that defines
the existence of connections (edges) between the nodes, thus with
entries A ij equal to zero if the nodes i and j are not linked by an
edge, or is different from zero if the nodes are connected. For a
graph weighted by w ij coefficients, the adjacency matrix reads
A ij ¼
w ij
if edge for nodes i and j exists
0
otherwise
(
ð3Þ
The zeros of the adjacency matrix should thus represent the
pairs of residues that do not exchange information (through correlated motions) and thus are not linked in the dynamical weighted
network. Sethi et al. [27] proposed to use a distance cutoff to
discriminate between residues that are linked or not in dynamical
weighted networks, i.e., suggested to include only edges that represent pairs of residues that are found at chemical distance (at least
one distance between heavy atoms is below the cutoff, varying in
the range of 3.5–5.5 A ˚ ) during the MD simulations, thus excluding
long-range correlations from the network of communications, see
Community Network Analysis of Allosteric Proteins
141
performed. The correlation coefficients can be computed for each
independent MD trajectory (or as a single computation of concatenated independent trajectories) and then averaged out. As mentioned above, the correlation coefficients (including r[x i ,x j ]) can be
computed using the GROMACS software, which requires the
topology file (containing all the necessary information to compute
the forces and velocities for the MD) to be in the GROMACS
format. This topology file format can be obtained from other
topology formats (e.g., from AMBER topology file) using free
software, such as ParmEd [46].
3 Methods
3.1 Dynamical
Weighted Network
The pair-correlations coefficients defined above measure how correlated are the motion of two residues recorded during a MD
simulation. When employing the mutual information measure,
i.e., the generalized r MI [x i ,x j ] coefficients, the pair-correlations
are directly related to the exchange of information between two
residues and can then be used to weight a graph that models the
information exchange within a protein according to its dynamics,
i.e., to build a dynamical weighted network. First, a protein graph
needs to be defined. Since the correlation coefficients are computed
using the C α atomic positions fluctuations, each node of the graph
is associated to each C α of the protein representing each amino acid
residue in the primary sequence, see Fig. 2a. To complete the graph,
one needs to build an adjacency matrix, A, i.e., a matrix that defines
the existence of connections (edges) between the nodes, thus with
entries A ij equal to zero if the nodes i and j are not linked by an
edge, or is different from zero if the nodes are connected. For a
graph weighted by w ij coefficients, the adjacency matrix reads
A ij ¼
w ij
if edge for nodes i and j exists
0
otherwise
(
ð3Þ
The zeros of the adjacency matrix should thus represent the
pairs of residues that do not exchange information (through correlated motions) and thus are not linked in the dynamical weighted
network. Sethi et al. [27] proposed to use a distance cutoff to
discriminate between residues that are linked or not in dynamical
weighted networks, i.e., suggested to include only edges that represent pairs of residues that are found at chemical distance (at least
one distance between heavy atoms is below the cutoff, varying in
the range of 3.5–5.5 A ˚ ) during the MD simulations, thus excluding
long-range correlations from the network of communications, see
Community Network Analysis of Allosteric Proteins
141
