2 Metagenome Analysis
53
The term binning describes the process in which sequence fragments of variable sizes are clustered into “bins”, where each of the bins most probably resembles
a single organism. Several binning approaches have been proposed over the last
years, the simplest just mapping occasionally occurring phylogenetic marker genes,
or all genes on the contigs, to taxonomic groups based on best BLAST-hits (Treusch
et al. 2004, Huson et al. 2007). More sophisticated approaches use intrinsic genomic
signatures based on oligonucleotide frequencies for fragment correlation or even
phylogenetic classification. Numerous studies have shown, that oligonucleotide
frequencies within DNA sequences exhibit species-specific patterns (Karlin et al.
1998); and for tetranucleotides it has even been demonstrated that their frequencies
carry an innate but weak phylogenetic signal (Pride et al. 2003). The currently available techniques can be classified into supervised (classification) and unsupervised
(clustering) methods (McHardy and Rigoutsos 2007). Here supervised means that
sequence fragments are classified based on their intrinsic DNA signatures to one
or many classes that have been modelled based on prior knowledge (e.g. all bacterial genomes). Recent examples for this approach are naïve Bayesian classifiers
(Sandberg et al. 2001) and Phylophytia (McHardy et al. 2007). The main advantage of these systems is that they provide a direct feedback about the phylogenetic
composition in the metagenome after presenting the metagenomic fragments to the
trained classifiers.
The main drawback of supervised classifiers is the availability of accurately classified sequence information for training. Optimal results are only obtained in case
sample specific classes e.g. trained on closely related sequenced genomes are available (Mavromatis et al. 2007, McHardy et al. 2007, Warnecke et al. 2007). This is
often problematic when working with environmental samples.
Unsupervised methods like TETRA (Teeling et al. 2004) or Self Organizing
Maps (SOM) (Abe et al. 2005, Abe et al. 2006) do not need a training set and produce sequence clusters independently of phylogenetic assignments. In a subsequent
mapping step phylogenetic marker genes, often present on at least on one fragment
in the clusters, are used to assign the sequence clusters to an organism. The clear
advantage of unsupervised methods is their ability to take a broad set of sequence
features, like GC content, read depth, into account, in addition to oligonucleotide
frequencies and distributions. These methods are able to assign metagenomic fragments to coherent organism bins, even if no prior information is available about
the prevailing organisms in a sample. The power of unsupervised correlation of
sequence fragments has been recently demonstrated for a consortium of microbes
involved in the anaerobic oxidation of methane (Krüger et al. 2003, Meyerdierks
et al. 2005), in single cell genomic analysis of Beggiotoa sp. (Mussmann et al.
2007) and in the reconstruction of the genomes of the four symbionts of the marine
oligochaete Olavius sp. (Woyke et al. 2006).
Nevertheless, both supervised and unsupervised methods are based on the statistical analysis of sequence data and therefore perform poorly if sequence fragments
are shorter than 5–8 kb. For assembled data this is not a major problem because currently the standard assembly methods fall short for contigs of less than 8 kb in length
(Mavromatis et al. 2007). A solution for the huge amount of single reads, which are
Précédent

- 68/410

Suivant