326
V. Mittard-Runte et al.
9.2.3 Common File Formats
Various different file formats have been developed in order to fulfil the requirements
of bioinformaticians and to provide a practical opportunity to share data between
different parties.
The FASTA format, originally specified by the National Center for Biotechnology Information (NCBI), has become the de-facto standard for the exchange
of raw sequence data. A FASTA file contains one or several sequences, with each
sequence being preceded by a single description line initiated with the “>” character. The subsequent lines contain the actual sequence data in human-readable
format; each nucleotide or amino acid is represented by a single letter following
the IUBMB/IUPAC standard code (http://www.chem.qmul.ac.uk/iupac/jcbn/).
Other commonly used file formats include the EMBL and GenBank format. Both
can contain more than simply the sequence data and are typically used to store e.g.
nucleotide or protein sequence data of a genome together with relevant annotation
information (see Section 9.3.5.1).
Several frameworks such as BioJava or BioPerl (Mangalam 2002) have been
developed in the bioinformatics field for the most common programming languages,
which offer an easy and convenient way to access and manipulate data contained in
these file formats.
9.3 DNA Sequence Analysis
In this section, we will present some bioinformatics tools available to marine biologists that want to use genomic approaches. Firstly we will consider eukaryotic EST
(Expressed Sequence Tag) processing as bacterial genome sequencing has already
been presented in the sequence data generation chapter. Then we will focus on gene
prediction, one of the first steps of the annotation of newly sequenced genomes. The
next step in working with genetic sequences is the prediction and analysis of the
functions and roles a gene or assembled EST might have. We use multiple methods
and information resources for sequence annotation such as InterPro. We will then
continue with an introduction to comparative genomics and functional classification, where the goal is to identify functions of particular regions of sequence data.
We will close this section by presenting the major public sequences databases and
other resources, which aim to provide a comprehensive coverage of sequences and
annotation available to the scientific community.
9.3.1 EST Processing
Many genome projects generate thousands of Expressed Sequence Tags (ESTs) or
shotgun reads. ESTs are generated by reverse transcribing mRNA into complementary DNA (cDNA), which is subsequently sequenced. ESTs provide a fast and
inexpensive manner to identify segments of DNA (a few hundred nucleotides) that
Précédent

- 337/410

Suivant