9 Genomic Techniques and How to Apply Them to Marine Questions
327
code for proteins and are expressed in a certain cell state or cell type. Several steps
are necessary to provide high quality sequences, as well as to get an overview of
their content. The analysis of EST sequences consists of the following steps:
• Pre-processing
• Clustering and assembly
• Functional analysis
(a) General Concept
EST pre-processing: After the sequencing step, the raw EST sequences are stored
as sequence trace files. These files are converted into FASTA files. The trimming
procedure removes low quality sequences at both ends of a read (quality clipping).
After quality clipping, the remaining parts of the cloning vector have to be
removed from the sequences, so that only the transcribed regions of the gene are left.
Therefore a BLAST search against a special database (containing vector sequences)
is performed and the vector sequences are removed (vector clipping).
EST clustering and assembly: To reduce redundancy, the EST sequences are grouped
(clustered) based on comparisons carried out at the DNA sequence level using a
clustering tool (e.g., TGICL – a clustering tool from The Institute for Genome
Research – TIGR) (Pertea et al. 2003). Different parameters have to be selected
for the clustering, including the length of the high-scoring segment pair (HSP), the
unmatched overhang length of the sequences, and a value for the identity of the
HSPs (see Fig. 9.2).
All EST sequences are searched for homology against all other ESTs using
MegaBLAST (Zhang et al. 2000) to determine which ESTs belong to each cluster.
The clusters are then assembled into Tentative Consensus sequences (TCs) (assembly). This procedure is carried out by a bioinformatics tool like CAP3 (Huang and
Madan 1999). It is possible to cluster and assemble ESTs from more than one
EST library together, so that genes occurring in different libraries (and therefore
expressed in different tissues/under different conditions) are assembled into one TC.
Figure 9.3 shows a schema of the clustering and assembling of ESTs to TCs. The
resulting TCs can be analysed functionally subsequently.
Fig. 9.2 This figure shows the overlap of two EST sequences. For the clustering procedure, the
high-scoring segment pair (HSP) length, the length of the unmatched overhang and the percentage
of the identity of the sequences have to be defined
327
code for proteins and are expressed in a certain cell state or cell type. Several steps
are necessary to provide high quality sequences, as well as to get an overview of
their content. The analysis of EST sequences consists of the following steps:
• Pre-processing
• Clustering and assembly
• Functional analysis
(a) General Concept
EST pre-processing: After the sequencing step, the raw EST sequences are stored
as sequence trace files. These files are converted into FASTA files. The trimming
procedure removes low quality sequences at both ends of a read (quality clipping).
After quality clipping, the remaining parts of the cloning vector have to be
removed from the sequences, so that only the transcribed regions of the gene are left.
Therefore a BLAST search against a special database (containing vector sequences)
is performed and the vector sequences are removed (vector clipping).
EST clustering and assembly: To reduce redundancy, the EST sequences are grouped
(clustered) based on comparisons carried out at the DNA sequence level using a
clustering tool (e.g., TGICL – a clustering tool from The Institute for Genome
Research – TIGR) (Pertea et al. 2003). Different parameters have to be selected
for the clustering, including the length of the high-scoring segment pair (HSP), the
unmatched overhang length of the sequences, and a value for the identity of the
HSPs (see Fig. 9.2).
All EST sequences are searched for homology against all other ESTs using
MegaBLAST (Zhang et al. 2000) to determine which ESTs belong to each cluster.
The clusters are then assembled into Tentative Consensus sequences (TCs) (assembly). This procedure is carried out by a bioinformatics tool like CAP3 (Huang and
Madan 1999). It is possible to cluster and assemble ESTs from more than one
EST library together, so that genes occurring in different libraries (and therefore
expressed in different tissues/under different conditions) are assembled into one TC.
Figure 9.3 shows a schema of the clustering and assembling of ESTs to TCs. The
resulting TCs can be analysed functionally subsequently.
Fig. 9.2 This figure shows the overlap of two EST sequences. For the clustering procedure, the
high-scoring segment pair (HSP) length, the length of the unmatched overhang and the percentage
of the identity of the sequences have to be defined
