2 Metagenome Analysis
57
the prediction of signal peptides with signalP (Bendtsen et al. 2004) and transmembrane helices prediction with TMHMM (Krogh et al. 2001). tRNAscan-SE (Lowe
and Eddy 1997) can be used to find and assign transfer RNAs within the sequence.
Once the calculations are finished, automatic annotation systems like Metanor,
provided by GenDB, or MicHanThi (Quast 2006) attempt to automatically generate annotations for all the predicted genes based on the observations returned
by the individual tools. This supports the manual annotation process by providing additional information for decision making. The MicHanThi system is currently
able to delineate and annotate hypothetical and conserved hypothetical genes in a
nearly quantitative manner. For genes with significant hits in primary or secondary
databases, such as UniProt or Pfam, the system is consistently able to assign the
correct functional category. In the subsequent manual annotation process, every
predicted gene should be investigated for significant hits to entries in the databases.
Starting with hits to Swiss-Prot and taking into account Pfam and InterPro results the
annotators have to integrate the information, read additional literature, and finally
assign a function to the genes. The history system implemented in GenDB tracks all
annotation changes and thus parallel annotations by different experts, or automatic
annotations systems, can be handled for every gene. After finishing the annotation
process, metabolic reconstructions can be performed to the extent that this is possible with the dataset in hand. The easiest way to do this is to automatically map
the EC-numbers to the corresponding KEGG pathway maps (Kanehisa et al. 2004)
provided by the GenDB system.
The rather slow web based visualisation system of GenDB has recently been
complemented by JCoast a software tool for data mining and comparison of
prokaryotic (meta)genomes (Richter et al. 2008). JCoast offers a flexible graphical user interface (GUI), as well as an application programming interface (API)
that facilitates direct back-end data access. The system offers individual genome,
cross-genome and metagenome analysis, and assists the biologist to explore large
and complex (meta)genomics datasets. The system can also work independently of
an existing GenDB installation as long as an appropriate database backend exists.
2.4.6 High Diversity Environments, Shallow Sequencing
and Short Read Technologies
Mastering the bioinformatic analysis of metagenomes becomes increasingly difficult as the environments studied increase in their diversity, especially if this factor
is combined with shallow sequencing. This has been nicely demonstrated by the
recently published Global Ocean Sampling campaigns (Venter et al. 2004, Seshadri
et al. 2007, Yooseph et al. 2007). Although billions of bases were obtained, only
two genomes could be reconstructed and reasonably large scaffolds could only be
assembled for the most dominant community members. Nevertheless, the incredible
repertoire of genes now available in our public databases has already stimulated a
broad range of follow up research aimed at investigating the diversity and function
57
the prediction of signal peptides with signalP (Bendtsen et al. 2004) and transmembrane helices prediction with TMHMM (Krogh et al. 2001). tRNAscan-SE (Lowe
and Eddy 1997) can be used to find and assign transfer RNAs within the sequence.
Once the calculations are finished, automatic annotation systems like Metanor,
provided by GenDB, or MicHanThi (Quast 2006) attempt to automatically generate annotations for all the predicted genes based on the observations returned
by the individual tools. This supports the manual annotation process by providing additional information for decision making. The MicHanThi system is currently
able to delineate and annotate hypothetical and conserved hypothetical genes in a
nearly quantitative manner. For genes with significant hits in primary or secondary
databases, such as UniProt or Pfam, the system is consistently able to assign the
correct functional category. In the subsequent manual annotation process, every
predicted gene should be investigated for significant hits to entries in the databases.
Starting with hits to Swiss-Prot and taking into account Pfam and InterPro results the
annotators have to integrate the information, read additional literature, and finally
assign a function to the genes. The history system implemented in GenDB tracks all
annotation changes and thus parallel annotations by different experts, or automatic
annotations systems, can be handled for every gene. After finishing the annotation
process, metabolic reconstructions can be performed to the extent that this is possible with the dataset in hand. The easiest way to do this is to automatically map
the EC-numbers to the corresponding KEGG pathway maps (Kanehisa et al. 2004)
provided by the GenDB system.
The rather slow web based visualisation system of GenDB has recently been
complemented by JCoast a software tool for data mining and comparison of
prokaryotic (meta)genomes (Richter et al. 2008). JCoast offers a flexible graphical user interface (GUI), as well as an application programming interface (API)
that facilitates direct back-end data access. The system offers individual genome,
cross-genome and metagenome analysis, and assists the biologist to explore large
and complex (meta)genomics datasets. The system can also work independently of
an existing GenDB installation as long as an appropriate database backend exists.
2.4.6 High Diversity Environments, Shallow Sequencing
and Short Read Technologies
Mastering the bioinformatic analysis of metagenomes becomes increasingly difficult as the environments studied increase in their diversity, especially if this factor
is combined with shallow sequencing. This has been nicely demonstrated by the
recently published Global Ocean Sampling campaigns (Venter et al. 2004, Seshadri
et al. 2007, Yooseph et al. 2007). Although billions of bases were obtained, only
two genomes could be reconstructed and reasonably large scaffolds could only be
assembled for the most dominant community members. Nevertheless, the incredible
repertoire of genes now available in our public databases has already stimulated a
broad range of follow up research aimed at investigating the diversity and function
