2 Metagenome Analysis
55
in the public databases. In general, errors can be introduced by inconsistencies in
functional assignments between and even within a single genome and by a simplistic procedure for assigning potential functions to the genes found. Unfortunately,
there is currently no “gold-standard” for consistent (meta)genome annotation available and there are no binding rules which have to be followed by all annotators
(Raes et al. 2007). To exploit the currently available data sources for functional predictions, comprehensive software systems are needed to store, analyse and visualise
data, and to support the decision process by providing information from various
sequence-based analysis tools. A “state-of-the-art” analysis pipeline for genomic
data should include gene finding and standard bioinformatic tools for similarity-,
pattern- and profile-based searches as well as prediction of signal peptides, transmembrane helices, transfer RNAs and other stable RNAs. In addition, the analysis
of global and local G+C content and skews as well as codon usage and other statistical parameters can be helpful in distinguishing coding from non-coding regions.
Adequate annotation systems should include automatic annotation of the protein
coding regions as well as user friendly annotation facilities for manual refinement via annotation jamborees (for more details see Stothard and Wishart 2006).
Finally, the reconstruction of metabolic pathways and networks helps to transform
the wealth of sequence information into biological knowledge.
2.4.4 Web Based Annotation Pipelines
Current, initiatives such as the Community Sequencing Program at the Joint
Genome Institute (JGI), the Microbial Genome Sequencing Project funded by the
Gordon and Betty Moore Foundation, or collaborations with Genoscope, have
enabled researchers worldwide to get their genome or metagenome of interest
sequenced. Moreover, these centres can fund a number of projects internally, solving
another important problem of cost. Unfortunately, when the sequences are recovered by the research laboratories, the processing of this data can pose a serious
problem due to a lack in expertise with installing, maintaining and/or implementing the software needed for genome annotation. Initially, bioinformatic support
is often provided by the sequencing centers through web-based systems, such as
the Integrated Microbial Genomes (IMG and IMG/M) system (Markowitz et al.
2006, 2008), Magnifying Genomes (Vallenet et al. 2006) or CAMERA (Seshadri
et al. 2007). Further online systems that accept raw (meta)genomic sequence data
for processing and provide web-based visualisation of results such as BASys and
PUMA2 (Van Domselaar et al. 2005, Maltsev et al. 2006) have been set up to
support functional assignments and metabolic reconstruction. Recently, the RAST
Server for rapid annotation using subsystem technology (Aziz et al. 2008) and
comparative genomics (Overbeek et al. 2005, Ye et al. 2005) has been released.
Besides gene prediction and annotation this system uses annotations to reconstruct
the metabolic networks that are functioning in the studied environment and then
makes all this data available for download. An experimental system specifically
55
in the public databases. In general, errors can be introduced by inconsistencies in
functional assignments between and even within a single genome and by a simplistic procedure for assigning potential functions to the genes found. Unfortunately,
there is currently no “gold-standard” for consistent (meta)genome annotation available and there are no binding rules which have to be followed by all annotators
(Raes et al. 2007). To exploit the currently available data sources for functional predictions, comprehensive software systems are needed to store, analyse and visualise
data, and to support the decision process by providing information from various
sequence-based analysis tools. A “state-of-the-art” analysis pipeline for genomic
data should include gene finding and standard bioinformatic tools for similarity-,
pattern- and profile-based searches as well as prediction of signal peptides, transmembrane helices, transfer RNAs and other stable RNAs. In addition, the analysis
of global and local G+C content and skews as well as codon usage and other statistical parameters can be helpful in distinguishing coding from non-coding regions.
Adequate annotation systems should include automatic annotation of the protein
coding regions as well as user friendly annotation facilities for manual refinement via annotation jamborees (for more details see Stothard and Wishart 2006).
Finally, the reconstruction of metabolic pathways and networks helps to transform
the wealth of sequence information into biological knowledge.
2.4.4 Web Based Annotation Pipelines
Current, initiatives such as the Community Sequencing Program at the Joint
Genome Institute (JGI), the Microbial Genome Sequencing Project funded by the
Gordon and Betty Moore Foundation, or collaborations with Genoscope, have
enabled researchers worldwide to get their genome or metagenome of interest
sequenced. Moreover, these centres can fund a number of projects internally, solving
another important problem of cost. Unfortunately, when the sequences are recovered by the research laboratories, the processing of this data can pose a serious
problem due to a lack in expertise with installing, maintaining and/or implementing the software needed for genome annotation. Initially, bioinformatic support
is often provided by the sequencing centers through web-based systems, such as
the Integrated Microbial Genomes (IMG and IMG/M) system (Markowitz et al.
2006, 2008), Magnifying Genomes (Vallenet et al. 2006) or CAMERA (Seshadri
et al. 2007). Further online systems that accept raw (meta)genomic sequence data
for processing and provide web-based visualisation of results such as BASys and
PUMA2 (Van Domselaar et al. 2005, Maltsev et al. 2006) have been set up to
support functional assignments and metabolic reconstruction. Recently, the RAST
Server for rapid annotation using subsystem technology (Aziz et al. 2008) and
comparative genomics (Overbeek et al. 2005, Ye et al. 2005) has been released.
Besides gene prediction and annotation this system uses annotations to reconstruct
the metabolic networks that are functioning in the studied environment and then
makes all this data available for download. An experimental system specifically
