9 Genomic Techniques and How to Apply Them to Marine Questions
323
technologies, together with a lower coverage using Sanger sequencing is often
optimal. This provides on the one hand sequence depth and on the other hand a
scaffold for assembly and finishing steps.
9.2 Data Management for Bioinformatics Applications
The ongoing developments of new sequencing technologies result in large amounts
of sequence data being generated. Together with the derived data that is created
from those sequences, processing of sequence data has become a task that not only
requires considerable computing resources, but also a large amount of data storage
capacity.
This section will provide an overview of the most important aspects that need to
be considered when planning genomic sequencing projects.
9.2.1 Data Modelling and Storage
A structured and well-organized data storage system is essential for data that are
frequently accessed. This involves not only the development of a formal data
description model, associated metadata and existing entity relationships that have
to be stored, but also an estimation of the amount of data that will be generated and
the allocation of required storage resources.
In a typical scenario, nucleotide or protein sequences can easily be saved to “flat”
files that contain only the sequence data arranged sequentially. Metadata like indices
are often created to allow faster access to individual records. But this approach,
while perfectly valid for this type of data, is generally unfeasible for other information like annotation data for genomes. When numerous relations exist within
the data, a relational database management system (Codd 1990) or another type of
structured storage is likely to be more appropriate.
For data that are frequently updated and changed, a centralized storage system
is preferred over keeping distributed copies of the data locally on each system.
However, such a centralized storage system might easily become a bottleneck when
stored data have to be accessed by a large number of systems (e.g. in a cluster
environment with many computers accessing a central sequence database). So it is
necessary to consider all the potential downsides and pitfalls that any solution might
have.
For data that are infrequently accessed, it might even be easier to recreate them
on demand and not to store them at all. Another option is to extract the relevant
details and to only store a subset of the original data for further processing.
Another important aspect that has to be considered is the availability of the data.
This includes not only implementing an access control scheme, which defines by
whom and how the data may be accessed, but also making a decision about the
Précédent

- 334/410

Suivant