66
2
Introduction
M. WILM et al.
Mass Spectrometry is efficiently used to identify proteins whose sequence had
been deposited in databases. For a protein whose complete sequence is known a
precise mass measurement of an ensemble of peptides produced by enzymatic
digestion of the protein can be sufficient to unambiguously retrieve its sequence
in a database (Jensen ON et al. 1996, Jensen ON et al. 1997, Jensen ON et al. 1996).
More specific and proven identifications can be obtained when the peptides are
subjected to tandem mass spectrometric investigations. Protein identities can be
revealed by reading out partial amino acid sequences from peptide fragment
spectra (sequence tag approach (Mann M 1996, Mann M and Wilm M 1994)) or
by correlating fragment spectra of pep tides to predicted spectra from peptides in
the database (Eng JK et al. 1994, Yates JR et al. 1995). Tandem MS based protein
identification can be done when only partial sequences are known, as for proteins reflected in EST databases (Mann M 1996). On picomole levels of protein it
is possible to reach a high throughput of dozens of identified proteins per day
(McCormack AL et al. 1997).
De novo protein sequencing with the aim to design oligonucleotide probes for
cloning the protein is a more difficult task. Cloning a protein with degenerate oligonucleotide primers can be time consuming even when the peptide sequences
are absolutely correct. Therefore, a de novo sequencing technique should not
only have a high sensitivity but a very high reliability as well. Two mass spectrometric techniques have been used to this end: MALD! high-energy collision
induced dissociation (Medzihradszky KF et al. 1996) and nano electrospray low
energy tandem mass spectrometry with C-terminal labeling of the peptides
(Hunt DF et al. 1986, Shevchenko A et al. 1997, Wilm M et al. 1996). More than 10
novel proteins could be cloned following peptide sequencing with the latter
approach (Bruyns E et al. 1998, Chen R-H et al. 1998, Lingner J et al. 1997,
McNagny KM et al. 1997, T. Kawata et al. 1997, Walczak H et al. 1997, Wilm M et
al. 1996).
The rational behind C-terminal peptide labeling is that C-terminal fragment
ions (y ions) can be unambiguously identified in the spectrum by the additional
mass of the label. Since trypsin cleaves after basic amino acids, the C-terminal
residue of a tryptic peptide is a charge retention site. Therefore, C-terminal fragments can be extracted by electrical fields from the collision zone of a tandem
MS mass spectrometer. As a consequence, fragment spectra of tryptic peptides
contain often long y ion series. When they can be identified via the incorporated
label long amino acid sequences can be retrieved.
C-terminal peptide labeling can be achieved by esterifying the pep tides with
methanol (Hunt DF et al. 1986) or by partial incorporation of 18 0 isotopes when
the protein is digested in 50 % 18 0 water (Shevchenko A et al. 1997). Methylation
increases the mass of y ions by 14 Da per free carboxyl group. This method was
employed for de novo sequencing when using a triple quadrupole mass spectrometer in our laboratory. Peptides were first analyzed in unmodified and then
in esterified form. Two conditions - precise amino acid spacings between adjacent y ions and a 14 Da mass shift per free carboxy group in the fragment of the
2
Introduction
M. WILM et al.
Mass Spectrometry is efficiently used to identify proteins whose sequence had
been deposited in databases. For a protein whose complete sequence is known a
precise mass measurement of an ensemble of peptides produced by enzymatic
digestion of the protein can be sufficient to unambiguously retrieve its sequence
in a database (Jensen ON et al. 1996, Jensen ON et al. 1997, Jensen ON et al. 1996).
More specific and proven identifications can be obtained when the peptides are
subjected to tandem mass spectrometric investigations. Protein identities can be
revealed by reading out partial amino acid sequences from peptide fragment
spectra (sequence tag approach (Mann M 1996, Mann M and Wilm M 1994)) or
by correlating fragment spectra of pep tides to predicted spectra from peptides in
the database (Eng JK et al. 1994, Yates JR et al. 1995). Tandem MS based protein
identification can be done when only partial sequences are known, as for proteins reflected in EST databases (Mann M 1996). On picomole levels of protein it
is possible to reach a high throughput of dozens of identified proteins per day
(McCormack AL et al. 1997).
De novo protein sequencing with the aim to design oligonucleotide probes for
cloning the protein is a more difficult task. Cloning a protein with degenerate oligonucleotide primers can be time consuming even when the peptide sequences
are absolutely correct. Therefore, a de novo sequencing technique should not
only have a high sensitivity but a very high reliability as well. Two mass spectrometric techniques have been used to this end: MALD! high-energy collision
induced dissociation (Medzihradszky KF et al. 1996) and nano electrospray low
energy tandem mass spectrometry with C-terminal labeling of the peptides
(Hunt DF et al. 1986, Shevchenko A et al. 1997, Wilm M et al. 1996). More than 10
novel proteins could be cloned following peptide sequencing with the latter
approach (Bruyns E et al. 1998, Chen R-H et al. 1998, Lingner J et al. 1997,
McNagny KM et al. 1997, T. Kawata et al. 1997, Walczak H et al. 1997, Wilm M et
al. 1996).
The rational behind C-terminal peptide labeling is that C-terminal fragment
ions (y ions) can be unambiguously identified in the spectrum by the additional
mass of the label. Since trypsin cleaves after basic amino acids, the C-terminal
residue of a tryptic peptide is a charge retention site. Therefore, C-terminal fragments can be extracted by electrical fields from the collision zone of a tandem
MS mass spectrometer. As a consequence, fragment spectra of tryptic peptides
contain often long y ion series. When they can be identified via the incorporated
label long amino acid sequences can be retrieved.
C-terminal peptide labeling can be achieved by esterifying the pep tides with
methanol (Hunt DF et al. 1986) or by partial incorporation of 18 0 isotopes when
the protein is digested in 50 % 18 0 water (Shevchenko A et al. 1997). Methylation
increases the mass of y ions by 14 Da per free carboxyl group. This method was
employed for de novo sequencing when using a triple quadrupole mass spectrometer in our laboratory. Peptides were first analyzed in unmodified and then
in esterified form. Two conditions - precise amino acid spacings between adjacent y ions and a 14 Da mass shift per free carboxy group in the fragment of the
