Part B | 10.7
300 Part B Tools and Methods in Marine Biotechnology
a 2 Gbp nonredundant reference database would take
3 days. Because BLAST computing time grows linearly with the reference database size, and the reference
databases double every year, a comparable growth in
the number of servers is needed. In other words, any
institution that aims to keep pace with the growth of
sequence data needs to double its budget each year.
Even when accounting for the 15% yearly increase
in CPU speed, a 1:7 factor budget growth is needed.
This simplified example does not yet incorporate concomitant costs incurred as a result of increases in
memory consumption, storage capacities, and network
bandwidth for data exchange, as well as power consumption and cooling. The issue becomes even more
severe when one considers that, unlike BLAST, several analysis tools have nonlinear computational time
consumption and the growth of metagenomic datasets
continues steadily.
The rapidly developing field of sequence-based
metagenomic analysis of marine microbiota promises
much for the future of marine biotechnology. While
sequence-based metagenomic approaches rely on comparing sequence data, which is obtained with sequences
deposited in databases, functional metagenomics focuses on screening DNA library clones directly for
a phenotype, i. e., the genes are recognizable by their
function rather than by their sequence. The power of
such an approach is that it does not require the genes
of interest to be recognizable by sequence analysis,
ensuring that this approach has the potential to directly identify entirely new classes of genes for both
known and, indeed, novel functions [10.25, 26]. Another advantage of such an approach is that while
sequence-based approaches can result in the incorrect
annotation of sequences with weak similarities to biochemically characterized gene products or of sequences
similar to gene products with multiple functions, the
results from a functional metagenomics approach are
unambiguous. Data from the Global Ocean Sampling
(GOS) expedition indicates that despite current largescale sequencing efforts the rate of discovery of new
protein families from the marine environment is linear, implying that marine microorganisms will continue
to be a source of novel enzymes in the foreseeable
future [10.27]. As many of these gene products are
entirely novel, their activity cannot be inferred from
comparison to known protein databases; thus a functional metagenomics approach has the ability to identify
novel genes on the basis of phenotypes which lend
themselves to high throughput screens Fig. 10.4.
10.7 Integrating Sequence and Contextual Data
From the time of field sampling to final step of sequence
analysis, a range of diverse data is produced among
different scientific communities, where individual researchers process the samples with different protocols in
different time frames. Often, the data is not deposited in
public resources at all, or the submitters have the choice
of up to a dozen repositories. Therefore, current molecular, environmental and diversity data is fragmented,
imprecise, or lost in lab books or proprietary private
archives. Moreover, environmental data, such as temperature, cannot be stored consistently in the current records
of sequences in the INSDC databases [10.28], nor can
a genome sequence be stored in databases dedicated to
environmental data. Unfortunately, even conceptually
simple contextual data-driven requests, such as Give me
the temperature at the sampling site of the microbial
isolate of interest or Give me all unknown genes sampled at temperatures greater than 80
ı C, are far from
trivial. The reason is as simple as it is profound.
An important source of contextual data is the data
taken in the field (on site), such as the geographic
and environmental origin of the sample, as well as information about subsequent processing to obtain the
DNA, and the sequencing itself. Given that each DNA
sequence from the marine environment is properly georeferenced, with geographic location, depth in the water
column or sediment, and time of sampling, the environmental context can be significantly complemented with
data from environmental databases. Only recently have
bioinformatic resources begun to invest in better management and integration of contextual data. CAMERA
hosts the fully georeferenced GOS data set. IMG/M
integrates rich details about the hosted genomes and
metagenomes. Megx.net is the first resource to provide
a comprehensive annotation of the environment of microbial genomes Kottmann et al., The Barcode of Life
initiative [10.29] successfully collaborated with INSDC
to store latitude and longitude in the public sequence
repositories. Classical ecological and conservation marine studies focused on species and communities; the
emphasis has now shifted to enhancing our understanding of the relationships among the various components
Précédent

- 340/1516

Suivant