3.6 De Novo Peptide
Identification
To determine the microbial species as well as protein composition
of the mucus samples, mass spectra are searched in two stages. First,
using a de novo search engine and followed by a targeted protein
database-dependent search. The purpose of the initial unbiased de
novo search is to identify all species in the samples in addition to the
host. De novo searches are preferred as the composition of the
microbial component is largely unknown and will not be hindered
in search time and accuracy in contrast to a database strategy against
all available species. When processed, the complete proteome databases for the species identified in the de novo search are combined
for the final targeted protein identification. Despite the fact that de
novo searches allow for the identification of many proteins, database-driven searches will currently outperform them and identify
more proteins. The combination of the two search strategies will
result in an unbiased assessment of the mucus metaproteome.
1. When required, convert mass spectrometry data files to peaklist
format for further data analysis (see Note 13).
2. Submit spectral data to the preferred de novo peptide search
engine. We recommend the use of the commercial software
PEAKS Studio v8.5 (Bioinformatic Solutions Inc.) for its performance and user-friendly interface. For alternative search
engines, see Note 14.
3. Use the following guidelines to set up the parameters for the de
novo searches: modifications: consider fixed only such as alkylating agents; search tolerance: set both precursor and fragment ion tolerance accurate for the MS instrument used (see
Note 15).
2 5
m
s
5 0
m
s
0
5000
10000
15000
1
2
3
4
5
6
0
50
100
Ascomycota
Bacteroidetes
Firmicutes
Mus musculus
Proteobacteria
Phylum_not_found
Actinobacteria
Streptophyta
Arthropoda
Basidiomycota
A L B
F C
G
B P
C
L C
A 1
K R T 1 9 Z G
1 6
A H
N
A K
S E R
P I N
A 3 K M
U
C
2
S E R
P I N
B 1 A T G
M
3
0
50
100
150
200
250
Spectral count de novo search
a
b
c
Fig. 2 De novo peptide search results overview. (a) Number of de novo mouse peptides identified using
increased ion accumulation time settings (n ¼ 6). (b) Top 10 phyla by a number of proteins identified by de
novo peptide sequencing (n ¼ 6). (c) Most abundant mouse proteins identified in the colonic mucus samples
based on spectral counts (n ¼ 6)
174
Sjoerd van der Post and Liisa Arike
Identification
To determine the microbial species as well as protein composition
of the mucus samples, mass spectra are searched in two stages. First,
using a de novo search engine and followed by a targeted protein
database-dependent search. The purpose of the initial unbiased de
novo search is to identify all species in the samples in addition to the
host. De novo searches are preferred as the composition of the
microbial component is largely unknown and will not be hindered
in search time and accuracy in contrast to a database strategy against
all available species. When processed, the complete proteome databases for the species identified in the de novo search are combined
for the final targeted protein identification. Despite the fact that de
novo searches allow for the identification of many proteins, database-driven searches will currently outperform them and identify
more proteins. The combination of the two search strategies will
result in an unbiased assessment of the mucus metaproteome.
1. When required, convert mass spectrometry data files to peaklist
format for further data analysis (see Note 13).
2. Submit spectral data to the preferred de novo peptide search
engine. We recommend the use of the commercial software
PEAKS Studio v8.5 (Bioinformatic Solutions Inc.) for its performance and user-friendly interface. For alternative search
engines, see Note 14.
3. Use the following guidelines to set up the parameters for the de
novo searches: modifications: consider fixed only such as alkylating agents; search tolerance: set both precursor and fragment ion tolerance accurate for the MS instrument used (see
Note 15).
2 5
m
s
5 0
m
s
0
5000
10000
15000
1
2
3
4
5
6
0
50
100
Ascomycota
Bacteroidetes
Firmicutes
Mus musculus
Proteobacteria
Phylum_not_found
Actinobacteria
Streptophyta
Arthropoda
Basidiomycota
A L B
F C
G
B P
C
L C
A 1
K R T 1 9 Z G
1 6
A H
N
A K
S E R
P I N
A 3 K M
U
C
2
S E R
P I N
B 1 A T G
M
3
0
50
100
150
200
250
Spectral count de novo search
a
b
c
Fig. 2 De novo peptide search results overview. (a) Number of de novo mouse peptides identified using
increased ion accumulation time settings (n ¼ 6). (b) Top 10 phyla by a number of proteins identified by de
novo peptide sequencing (n ¼ 6). (c) Most abundant mouse proteins identified in the colonic mucus samples
based on spectral counts (n ¼ 6)
174
Sjoerd van der Post and Liisa Arike
