60
). R. YATES et al.
5
Post Search Data Analysis
Autoquest is a script that runs a collection of programs. Included in Autoquest is
a program to process search results into a summarized and easily readable format. In addition to preprocessing data files, performing searches, and running a
summary, the following data review process occurs: The first step fIlters the summary fIle to a single entry per tandem mass spectrum. This process entails selecting between the better search result of the +2 and +3 charge state of the same
spectrum. The initial criteria for making this selection between charge states is
the raw cross correlation score. If the correlation score of one charge state is significantly larger than the other then that result is selected. Otherwise, the determination is based on a combination of the preliminary score rankings, the raw
correlation scores, and the peptide sequences and how well they follow the cleavage rules of trypsin. After the first step is completed, another program generates
an HTML formatted fIle containing search statistics and displays the proteins and
peptides identified. The program can process the results of more then one search.
The result lists each protein and the peptides identified as well as the number of
times the peptide was identified and its relative score. The user has control over
the scoring criteria (cross correlation scoring ranges and requiring peptides to be
tryptic within each scoring range) and can interactively fIlter and view the results
via a web browser. What makes this program particularly useful is the ability to
analyze multiple SEQUEST runs together so that identifications and statistics can
be tracked across several analyses.
6
Protein Mixture Analysis
Protein mixture analysis is made possible by the capability of tandem mass spectrometers to characterize a single component in a mixture of many. To perform
protein mixture analysis, protein mixtures are first digested with a protease such
as trypsin. The complicated peptide mixture is then subjected to separation online followed by automated tandem mass spectrometry data acquisition. There
are several fundamental reasons for adopting this approach. First, the diverse
chemistry inherent to proteins results in a wide range of solubility. Proteins most
frequently missing in a 2-D SDS-PAGE analysis are those of limited solubility
(Garrels et al. 1997). We expect each protein will produce at least one, if not
many, soluble peptides within the scan range of the mass spectrometer. Second,
the technology to sequence pep tides using tandem mass spectrometry is well
established and robust (Hunt et al. 1992; Hunt et al. 1986). The limit of detection
for protein identification is more dependent on the mass spectrometer then the
separation, visualization, and manipulation techniques necessary to isolate a
homogeneous protein. Lastly, the number of steps involved in sample handling is
reduced. Manipulation is performed with complex mixtures of pep tides or proteins at stages where sample losses would be minimal because the large amount
of peptide/protein present acts as a carrier. Most significant sample loss generally
occurs during attempts to manipulate small quantities of homogeneous material.
). R. YATES et al.
5
Post Search Data Analysis
Autoquest is a script that runs a collection of programs. Included in Autoquest is
a program to process search results into a summarized and easily readable format. In addition to preprocessing data files, performing searches, and running a
summary, the following data review process occurs: The first step fIlters the summary fIle to a single entry per tandem mass spectrum. This process entails selecting between the better search result of the +2 and +3 charge state of the same
spectrum. The initial criteria for making this selection between charge states is
the raw cross correlation score. If the correlation score of one charge state is significantly larger than the other then that result is selected. Otherwise, the determination is based on a combination of the preliminary score rankings, the raw
correlation scores, and the peptide sequences and how well they follow the cleavage rules of trypsin. After the first step is completed, another program generates
an HTML formatted fIle containing search statistics and displays the proteins and
peptides identified. The program can process the results of more then one search.
The result lists each protein and the peptides identified as well as the number of
times the peptide was identified and its relative score. The user has control over
the scoring criteria (cross correlation scoring ranges and requiring peptides to be
tryptic within each scoring range) and can interactively fIlter and view the results
via a web browser. What makes this program particularly useful is the ability to
analyze multiple SEQUEST runs together so that identifications and statistics can
be tracked across several analyses.
6
Protein Mixture Analysis
Protein mixture analysis is made possible by the capability of tandem mass spectrometers to characterize a single component in a mixture of many. To perform
protein mixture analysis, protein mixtures are first digested with a protease such
as trypsin. The complicated peptide mixture is then subjected to separation online followed by automated tandem mass spectrometry data acquisition. There
are several fundamental reasons for adopting this approach. First, the diverse
chemistry inherent to proteins results in a wide range of solubility. Proteins most
frequently missing in a 2-D SDS-PAGE analysis are those of limited solubility
(Garrels et al. 1997). We expect each protein will produce at least one, if not
many, soluble peptides within the scan range of the mass spectrometer. Second,
the technology to sequence pep tides using tandem mass spectrometry is well
established and robust (Hunt et al. 1992; Hunt et al. 1986). The limit of detection
for protein identification is more dependent on the mass spectrometer then the
separation, visualization, and manipulation techniques necessary to isolate a
homogeneous protein. Lastly, the number of steps involved in sample handling is
reduced. Manipulation is performed with complex mixtures of pep tides or proteins at stages where sample losses would be minimal because the large amount
of peptide/protein present acts as a carrier. Most significant sample loss generally
occurs during attempts to manipulate small quantities of homogeneous material.
