9 Genomic Techniques and How to Apply Them to Marine Questions
329
expressed genes. The sequence and annotation data can be exported in various file
formats. SAMS uses an object-oriented backend and a relational database management system MySQL, connected by an object relational mapping developed by the
Bioinformatics Resource Facility at Bielefeld University.
Users need to have a username and password. Access to the project can be
restricted to any number of users, so that the data are kept confidential.
SAMS can be accessed via a web browser (http://www.cebitec.uni-bielefeld.
de/groups/brf/software/sams/). Passwords can be requested via the login page.
Using SAMS for the analysis of EST sequences is possible in an easy and comfortable way. Firstly, a library needs to be selected for the sequences. The user
creates or specifies a library in which to store the sequence data. Then the sequences
are imported using the import function. This involves several consecutive steps.
First, the user selects the locally stored file to upload to the SAMS server. Several
formats are supported but raw data files are recommended. In the second step, raw
chromatogram files or trace files are uploaded and the program PHRED is used for
the quality clipping. The user can select the quality that is required (e.g., PHRED 13
gives a base-calling accuracy of 95%, the probability that the base called is wrong
is 1 in 20; PHRED 20 gives you 99% and a probability of 1 in 100). Finally, if the
sequences still contain remaining vector sequences, the user has the possibility to
remove them using the integrated vector clipping option.
Note that sequences shorter than 50 bp are removed after these procedures. Once
the import is finished, the sequences will be listed as ESTs in the SAMS library. The
user can then start a clustering procedure to reduce redundancy in the sequence data.
For that, standard parameters are suggested, but can be changed if required. The
standard clustering parameter set is the TIGR default parameter set (HSP length:
40, identity: 0.95, unmatched overhang length: 20). SAMS uses a clustering component, which is a TGICL-like implementation of the TIGR clustering approach.
After clustering, the sequence data is assembled into Tentative Consensus sequences
(TCs) using the program CAP3. In SAMS it is possible to cluster and assemble
the same sequence data more than once to test different parameter settings. When
the user is satisfied with the results of the clustering, the assembly task needs to
be performed in order to visualise the TCs and singletons, which are stored in the
database.
Functional analysis in SAMS: The goal of this part of the analysis is to achieve
the annotation of TCs. An automatic annotation pipeline called Metanor (Goesmann
et al. 2005) can be used for this, after the assembly procedure has been completed. This pipeline consists of several bioinformatics tools that are run for each
TC, EST and singleton. The tools use BLAST (Basic Local Alignment Search
Tool) to search for homology between the sequences and sequences contained in
different databases, like the NT (Non-redundant nucleotide database from NCBI),
NR (Non-redundant protein database from NCBI), KEGG, KOG/COG (clusters of
euKaryotic Orthologous Groups or Clusters of Orthologous Groups, see Section
9.3.4.4), and SwissProt databases (see Section 9.3.5.2). In addition, InterProScan
(see Section 9.3.3.3) is performed for all six reading frames. The results are stored
as “observations” in the database for further usage.
329
expressed genes. The sequence and annotation data can be exported in various file
formats. SAMS uses an object-oriented backend and a relational database management system MySQL, connected by an object relational mapping developed by the
Bioinformatics Resource Facility at Bielefeld University.
Users need to have a username and password. Access to the project can be
restricted to any number of users, so that the data are kept confidential.
SAMS can be accessed via a web browser (http://www.cebitec.uni-bielefeld.
de/groups/brf/software/sams/). Passwords can be requested via the login page.
Using SAMS for the analysis of EST sequences is possible in an easy and comfortable way. Firstly, a library needs to be selected for the sequences. The user
creates or specifies a library in which to store the sequence data. Then the sequences
are imported using the import function. This involves several consecutive steps.
First, the user selects the locally stored file to upload to the SAMS server. Several
formats are supported but raw data files are recommended. In the second step, raw
chromatogram files or trace files are uploaded and the program PHRED is used for
the quality clipping. The user can select the quality that is required (e.g., PHRED 13
gives a base-calling accuracy of 95%, the probability that the base called is wrong
is 1 in 20; PHRED 20 gives you 99% and a probability of 1 in 100). Finally, if the
sequences still contain remaining vector sequences, the user has the possibility to
remove them using the integrated vector clipping option.
Note that sequences shorter than 50 bp are removed after these procedures. Once
the import is finished, the sequences will be listed as ESTs in the SAMS library. The
user can then start a clustering procedure to reduce redundancy in the sequence data.
For that, standard parameters are suggested, but can be changed if required. The
standard clustering parameter set is the TIGR default parameter set (HSP length:
40, identity: 0.95, unmatched overhang length: 20). SAMS uses a clustering component, which is a TGICL-like implementation of the TIGR clustering approach.
After clustering, the sequence data is assembled into Tentative Consensus sequences
(TCs) using the program CAP3. In SAMS it is possible to cluster and assemble
the same sequence data more than once to test different parameter settings. When
the user is satisfied with the results of the clustering, the assembly task needs to
be performed in order to visualise the TCs and singletons, which are stored in the
database.
Functional analysis in SAMS: The goal of this part of the analysis is to achieve
the annotation of TCs. An automatic annotation pipeline called Metanor (Goesmann
et al. 2005) can be used for this, after the assembly procedure has been completed. This pipeline consists of several bioinformatics tools that are run for each
TC, EST and singleton. The tools use BLAST (Basic Local Alignment Search
Tool) to search for homology between the sequences and sequences contained in
different databases, like the NT (Non-redundant nucleotide database from NCBI),
NR (Non-redundant protein database from NCBI), KEGG, KOG/COG (clusters of
euKaryotic Orthologous Groups or Clusters of Orthologous Groups, see Section
9.3.4.4), and SwissProt databases (see Section 9.3.5.2). In addition, InterProScan
(see Section 9.3.3.3) is performed for all six reading frames. The results are stored
as “observations” in the database for further usage.
