8
D. R. GOODLETT et al.
using appropriate database search tools. A particularly powerful database search
tool is SEQUEST because it uses highly informative peptide cm spectra to automatically search sequence databases (Ducret et al. 1998). In principle and in
practice (Santos, et al. 1999), an unknown protein which has never been
sequenced can be identified using a single good quality cm spectrum.
Most proteins contain at least a few basic residues such as lysine and arginine
that trypsin will recognize and cleave on the C-terminal side of the amide bond.
The products of this specific digestion are peptides with protonation sites at the
C-terminii (i.e. lysine or arginine) that can, in the absence of interference by
other amino acids, generate a complete "ladder" series of y-ions (Roepstorff,
1984) in the tandem mass spectrum (Hunt et al. 1986). For this reason manual
confirmation of the sequence provided by available routines like SEQUEST is also
made easier when trypsin is used. Note that it is always important to manually
check the answers provided by computer search routines because scores are
based on a set of rigid rules rather than knowledge of protein chemistry. Other
proteolytic enzymes cannot be assumed to produce such a beneficial effect, but
are still useful when trypsin fails to produce peptides in the 1000-2000 Da range.
5
Challenges of Low Abundance Proteins
In yeast a total of approximately 6000 genes are present (Goffeau et al. 1996). To
determine what fraction of these genes are visualized by 2DE when a total cell
lysate is separated by 2DE, a recent study in our laboratory examined the codon
bias value for proteins identified after separation by 2DE. The codon bias value is
a calculated measure of the degree of redundancy of triplet DNA codons used to
produce each amino acid in a given gene sequence. It is a useful indicator of the
level of gene products in a cell (Garrels, et al. 1997). It is assumed that codon bias
correlates with the protein copy number per cell, with high codon bias indicating
a high level of protein expression. Codon bias values calculated for the entire
yeast genome showed that most proteins have a codon bias of <0.2 (Fig. 1.3A)
and are therefore of low abundance or low copy number per cell. It was shown,
however, that the majority of yeast proteins that could be readily identified from
silver stained 2DE gels (Gygi et al. 1999) had a codon bias value >0.2 (Fig. l.3B).
Therefore the proteins visible by our most sensitive, generic staining technique,
are the most abundant proteins in yeast cells. The proteins not visible by silver
staining are likely to be the very proteins of interest as targets for small molecule
therapeutics. To make these proteins available for analysis will require new methods to be devised to enrich celllysates for these proteins.
As mentioned earlier the best 2DE gels can separate in excess of several thousand proteins including regulatory proteins such as kinases, phosphatases, and
growth factors. Very specific human cell types are thought to express around
50,000 to 100,000 proteins. Therefore, only a small fraction of the total protein
present can be visualized by 2DE. Limitations in loading capacity and resolution
of 2DE can be circumvented by fractionation prior to loading. This sort of purification could be as simple as an ammonium sulfate fractionation or bulk cation
exchange prior to loading. The problem raised by an examination of codon bias
D. R. GOODLETT et al.
using appropriate database search tools. A particularly powerful database search
tool is SEQUEST because it uses highly informative peptide cm spectra to automatically search sequence databases (Ducret et al. 1998). In principle and in
practice (Santos, et al. 1999), an unknown protein which has never been
sequenced can be identified using a single good quality cm spectrum.
Most proteins contain at least a few basic residues such as lysine and arginine
that trypsin will recognize and cleave on the C-terminal side of the amide bond.
The products of this specific digestion are peptides with protonation sites at the
C-terminii (i.e. lysine or arginine) that can, in the absence of interference by
other amino acids, generate a complete "ladder" series of y-ions (Roepstorff,
1984) in the tandem mass spectrum (Hunt et al. 1986). For this reason manual
confirmation of the sequence provided by available routines like SEQUEST is also
made easier when trypsin is used. Note that it is always important to manually
check the answers provided by computer search routines because scores are
based on a set of rigid rules rather than knowledge of protein chemistry. Other
proteolytic enzymes cannot be assumed to produce such a beneficial effect, but
are still useful when trypsin fails to produce peptides in the 1000-2000 Da range.
5
Challenges of Low Abundance Proteins
In yeast a total of approximately 6000 genes are present (Goffeau et al. 1996). To
determine what fraction of these genes are visualized by 2DE when a total cell
lysate is separated by 2DE, a recent study in our laboratory examined the codon
bias value for proteins identified after separation by 2DE. The codon bias value is
a calculated measure of the degree of redundancy of triplet DNA codons used to
produce each amino acid in a given gene sequence. It is a useful indicator of the
level of gene products in a cell (Garrels, et al. 1997). It is assumed that codon bias
correlates with the protein copy number per cell, with high codon bias indicating
a high level of protein expression. Codon bias values calculated for the entire
yeast genome showed that most proteins have a codon bias of <0.2 (Fig. 1.3A)
and are therefore of low abundance or low copy number per cell. It was shown,
however, that the majority of yeast proteins that could be readily identified from
silver stained 2DE gels (Gygi et al. 1999) had a codon bias value >0.2 (Fig. l.3B).
Therefore the proteins visible by our most sensitive, generic staining technique,
are the most abundant proteins in yeast cells. The proteins not visible by silver
staining are likely to be the very proteins of interest as targets for small molecule
therapeutics. To make these proteins available for analysis will require new methods to be devised to enrich celllysates for these proteins.
As mentioned earlier the best 2DE gels can separate in excess of several thousand proteins including regulatory proteins such as kinases, phosphatases, and
growth factors. Very specific human cell types are thought to express around
50,000 to 100,000 proteins. Therefore, only a small fraction of the total protein
present can be visualized by 2DE. Limitations in loading capacity and resolution
of 2DE can be circumvented by fractionation prior to loading. This sort of purification could be as simple as an ammonium sulfate fractionation or bulk cation
exchange prior to loading. The problem raised by an examination of codon bias
