Direct Analysis of Protein Complexes
59
normal protein. The remaining spectra were searched through a hemoglobin
database. As shown in Fig. 4.2 a majority of the remaining tandem mass spectra
match to sequences corresponding to the mutated site of E ~ V, the mutation
responsible for sickle cell anemia.
4
Database Searching for Protein Identification
The most important element of protein identification in mixtures is the ability to
search sequence databases using tandem mass spectra. Protein identification
based on tandem mass spectrometry data was introduced in 1994 by Eng et al.
(Eng et al. 1994) The process uses both the molecular weight of the peptide and
the fragmentation pattern produced in collision induced dissociation. A peptide's
molecular weight is indicative of amino acid composition while the fragmentation pattern is directly related to the amino acid sequence of the peptide. In the
cm process peptides fragment primarily at amide bonds thus producing patterns that are indicative of the sequence of amino acids. Eng et al utilizes the
molecular weight of the peptide as a first criterion to identify amino acid
sequences in the database. Once candidate sequences are found, the peptide's
sequence ions are predicted. The ions predicted by the sequence are compared to
the ions present in the tandem mass spectrum. The intensities of sequence ions
present in the spectrum are summed. A closeness-of-fit measure is calculated for
each sequence within the molecular weight tolerance. These scores are ranked
and the top 500 sequences are then compared to the tandem mass spectrum
using correlation analysis. This analysis is performed by creating a model tandem mass spectrum for each sequence and comparing this model spectrum to
the experimental tandem mass spectrum using a cross-correlation function. The
better the fit between the two "spectra" the greater the value of the correlation.
A typical LC/MS/MS analysis using data-dependent acquisition will create
hundreds or more tandem mass spectra. To increase search speed through a
database the analysis can be multiplexed using multi-processor computers. The
analysis is well suited for this approach since each spectrum can be sent to a different processor for analysis resulting in an almost linear increase in speed. We
have created a version of SEQUEST for multiplexing the analysis of tandem mass
spectra across multiple computers. The software uses the Parallel Virtual
Machine (PVM) message passing software to create a virtual parallel processing
computer. A computer system has been built based on the Beowolf configuration
using 12 alpha based 533 MHz processors running the Linux operating system. At
the start of the analysis each computer processor is sent a spectrum. As each processor completes a search, the results are transmitted back to the master computer with a request to send another spectrum. The total processing time of 500
spectra through the non-redundant protein sequence databases (313,000 entries)
is 1 minute and 37 seconds. The Unigene EST sequence (53,000 entries) database
can be searched with a 3-frame translation in 1 minute and 16 seconds. These
improved search speeds provide a greater level of efficiency for the analysis of
protein mixtures and offer the possibility of real-time data analysis.
Précédent

- 69/371

Suivant