engineer proteins (and ligands) and optimize molecular properties
such as binding affinity and binding specificity [1–4]. It can explore
thousands of polypeptide sequences in 1–2 days using a desktop
computer. PDZ complexes can also be studied with mediumthroughput approaches that combine molecular dynamics
(MD) for conformational exploration with a simplified free energy
function for binding affinities. These approaches [5–15] can
explore a few dozen sequences in a few days using a few GPU
computers or a small Linux cluster.
Here, we outline both approaches and illustrate them by our
recent work on the Tiam1 PDZ domain. In a typical design project,
one might run a high-throughput calculation first, then characterize the top sequences with the medium-throughput approach. The
latter step might lead to additional high-throughput calculations,
perhaps involving additional mutating positions, and so on, in a
series of design cycles that would hopefully include experimental
steps. The illustrations below include high-throughput CPD protocols that can be run with our Proteus software [16, 17]. Our
medium-throughput approach uses a free energy function that
includes a Poisson–Boltzmann (PB) contribution, a surface area
(SA) contribution, and a van der Waals (vdW) contribution. It can
be run with standard MD and PB software.
1.1 General Issues
When engineering several amino acid positions, the mutation space
grows exponentially with the number of positions. If just three
positions are explored combinatorially, there are 8000 possible
sequences (more, if ncAAs are allowed). Each sequence variant
can occupy a vast number of conformations. Even if we use a
model where the protein backbone is held fixed, there are
thousands of rotamer combinations for the side chains at the
three mutating side positions, and millions if we consider the
whole binding interface. Another difficulty is that to engineer
PDZ–peptide binding, we should consider both the bound and
unbound states of each partner. If we engineer the protein, its
stability should be maintained, so that the unfolded state will also
play a role. If we care about specificity, we may have to design
against off-target binding, which would involve additional proteins
and states.
To address all these issues, major approximations are needed for
both conformational exploration and scoring. The first common
approximation is to use a molecular mechanics description of the
protein and peptide. Although some CPD tools like Rosetta use
knowledge-based energy functions [18, 19], many studies have
shown that molecular mechanics is a good approximation for
many design problems. The second common approximation is to
model solvent implicitly, as in continuum electrostatics. Formally,
this can be viewed as an averaging operation over solvent configurations [20]. This leads to the concept of a potential of mean force
238
Nicolas Panel et al.
such as binding affinity and binding specificity [1–4]. It can explore
thousands of polypeptide sequences in 1–2 days using a desktop
computer. PDZ complexes can also be studied with mediumthroughput approaches that combine molecular dynamics
(MD) for conformational exploration with a simplified free energy
function for binding affinities. These approaches [5–15] can
explore a few dozen sequences in a few days using a few GPU
computers or a small Linux cluster.
Here, we outline both approaches and illustrate them by our
recent work on the Tiam1 PDZ domain. In a typical design project,
one might run a high-throughput calculation first, then characterize the top sequences with the medium-throughput approach. The
latter step might lead to additional high-throughput calculations,
perhaps involving additional mutating positions, and so on, in a
series of design cycles that would hopefully include experimental
steps. The illustrations below include high-throughput CPD protocols that can be run with our Proteus software [16, 17]. Our
medium-throughput approach uses a free energy function that
includes a Poisson–Boltzmann (PB) contribution, a surface area
(SA) contribution, and a van der Waals (vdW) contribution. It can
be run with standard MD and PB software.
1.1 General Issues
When engineering several amino acid positions, the mutation space
grows exponentially with the number of positions. If just three
positions are explored combinatorially, there are 8000 possible
sequences (more, if ncAAs are allowed). Each sequence variant
can occupy a vast number of conformations. Even if we use a
model where the protein backbone is held fixed, there are
thousands of rotamer combinations for the side chains at the
three mutating side positions, and millions if we consider the
whole binding interface. Another difficulty is that to engineer
PDZ–peptide binding, we should consider both the bound and
unbound states of each partner. If we engineer the protein, its
stability should be maintained, so that the unfolded state will also
play a role. If we care about specificity, we may have to design
against off-target binding, which would involve additional proteins
and states.
To address all these issues, major approximations are needed for
both conformational exploration and scoring. The first common
approximation is to use a molecular mechanics description of the
protein and peptide. Although some CPD tools like Rosetta use
knowledge-based energy functions [18, 19], many studies have
shown that molecular mechanics is a good approximation for
many design problems. The second common approximation is to
model solvent implicitly, as in continuum electrostatics. Formally,
this can be viewed as an averaging operation over solvent configurations [20]. This leads to the concept of a potential of mean force
238
Nicolas Panel et al.
