ΔG s À ΔG r ¼ ÀkT ln
p
0
s
p 0
r
þ kT ln
p s
p r
ð3Þ
Extraction of the populations from the MC simulations and calculation of the free energy is done by Proteus with a python script. For
sequences of interest, such as the tightest binders, 3D structure
models can be computed from the rotamer information with a bash
script.
2.2.4 Application
to the Tiam1–Sdc1
Complex
An application to the Tiam1–Sdc1 complex was reported recently
[22]. Four of the last five peptide positions were allowed to mutate
into all types except Gly or Pro. The position numbers were À4 to
À1, following the usual “backward” convention for PDZ binders.
The C-terminal position (position 0) was kept as in the wild-type
peptide (Ala), because the Sdc1 backbone arrangement does not
allow large side chains at this position. There were 104,976 possible
sequences. Thanks to the adaptive method, a large fraction were
sampled. For nine variants, relative binding free energies were
available from experiment or high-level, alchemical MD free energy
simulations. Excluding one large error, the mean unsigned errors
from eight variants was 0.8 kcal/mol. Figure 2 shows the sequences
sampled, in the form of a sequence logo, with populations given by
their relative binding free energies. The logo is compared to one
that represents an experimental library of Tiam1-binding peptides
[43]. Positions P 0 and P À1 are the most important for PDZ binding
specificity. Position P 0 occupies three main types experimentally, C,
A, and F, but was held fixed during the simulations. Of the four
positions allowed to mutate, P À1 , P À3 , and P À4 are highly variable
in both the MC and the experimental logos. Of the top ten MC
types at these positions, 7 or 8 are present in the experimental logo,
and vice versa, with somewhat different occupancies. Position P À2
is more conserved, both experimentally and in the simulations. Of
the top four experimental types, Y, F, M, T, all but T are in the top
five MC types. While T has less than 1% occupancy in the MC
sequences, the chemically similar types A, C, and S are highly
populated. Overall, the two logos are in reasonable agreement.
3 A Medium-Throughput Design Approach
After high-throughput design, one can use a more costly model to
characterize a few dozen of the top CPD candidates with increased
accuracy [44]. The model uses MD with explicit solvent to sample
conformations, then scores them with a free energy function where
solvent is modeled implicitly. MD is done for the PDZ–peptide
complex, typically for 80–100 ns. Several hundred snapshots are
taken from the trajectory. The unbound state is modeled by using
the same snapshots and simply separating the two partners. We refer
to this as a single-trajectory approach. A two-trajectory variant that
244
Nicolas Panel et al.
p
0
s
p 0
r
þ kT ln
p s
p r
ð3Þ
Extraction of the populations from the MC simulations and calculation of the free energy is done by Proteus with a python script. For
sequences of interest, such as the tightest binders, 3D structure
models can be computed from the rotamer information with a bash
script.
2.2.4 Application
to the Tiam1–Sdc1
Complex
An application to the Tiam1–Sdc1 complex was reported recently
[22]. Four of the last five peptide positions were allowed to mutate
into all types except Gly or Pro. The position numbers were À4 to
À1, following the usual “backward” convention for PDZ binders.
The C-terminal position (position 0) was kept as in the wild-type
peptide (Ala), because the Sdc1 backbone arrangement does not
allow large side chains at this position. There were 104,976 possible
sequences. Thanks to the adaptive method, a large fraction were
sampled. For nine variants, relative binding free energies were
available from experiment or high-level, alchemical MD free energy
simulations. Excluding one large error, the mean unsigned errors
from eight variants was 0.8 kcal/mol. Figure 2 shows the sequences
sampled, in the form of a sequence logo, with populations given by
their relative binding free energies. The logo is compared to one
that represents an experimental library of Tiam1-binding peptides
[43]. Positions P 0 and P À1 are the most important for PDZ binding
specificity. Position P 0 occupies three main types experimentally, C,
A, and F, but was held fixed during the simulations. Of the four
positions allowed to mutate, P À1 , P À3 , and P À4 are highly variable
in both the MC and the experimental logos. Of the top ten MC
types at these positions, 7 or 8 are present in the experimental logo,
and vice versa, with somewhat different occupancies. Position P À2
is more conserved, both experimentally and in the simulations. Of
the top four experimental types, Y, F, M, T, all but T are in the top
five MC types. While T has less than 1% occupancy in the MC
sequences, the chemically similar types A, C, and S are highly
populated. Overall, the two logos are in reasonable agreement.
3 A Medium-Throughput Design Approach
After high-throughput design, one can use a more costly model to
characterize a few dozen of the top CPD candidates with increased
accuracy [44]. The model uses MD with explicit solvent to sample
conformations, then scores them with a free energy function where
solvent is modeled implicitly. MD is done for the PDZ–peptide
complex, typically for 80–100 ns. Several hundred snapshots are
taken from the trajectory. The unbound state is modeled by using
the same snapshots and simply separating the two partners. We refer
to this as a single-trajectory approach. A two-trajectory variant that
244
Nicolas Panel et al.
