358
V. Mittard-Runte et al.
long chains of tags. In this classical protocol, the tags are then cloned into a plasmid vector and their sequence is obtained using Sanger sequencing. The specific
cleavage site remains in the DNA-vector and can be used as a separator.
If genomic sequences of the organism are known, either in the form of ESTs or
as whole genome, the tags obtained from sequencing can be mapped onto the longer
sequences, thereby enabling a quantitative analysis of differential gene-expression.
The number of similar tags found can be assumed to correlate (although not linearly
due to the involvement of PCR-amplification) with the original amount of cDNA
fragments from which these tags are uniquely derived.
If no sequence information is available for the source organism, the sequence
tags can still be analysed quantitatively, as the tag sequence is an (almost unique)
identifier for the corresponding mRNA. However, mapping of short tags of 11 bp to
available databases for sequence comparison will yield a large proportion of false
positive hits.
An initial problem of the SAGE protocol was the high amount of source poly-A
RNA required (approximately 2.5–5 μg). Therefore, Datson et al. (1999) developed
MicroSAGE which uses strepatividin coated tubes instead of beads. The method
requires only 1 ng of RNA, which is approximately the equivalent of 100,000
mammalian cells
Another drawback of the initial SAGE approach was the short tag-length. Several
enhancements of the technique have been published which provide an improved taglength by using different restriction enzymes. Tag lengths of 20 bp (LongSAGE)
(Saha et al. 2002), and 26 bp (SuperSAGE) (Matsumura et al. 2003), have been
reported. These provide increasingly significant database hits.
Another improvement has been the development of gene identification methods
based on paired-end di-tags (GIS-PETs) (Ng et al. 2005). In contrast to SAGE, PETs
are derived from the 3 and 5 signatures of mRNA. For the 3 end, primers specific to the capping structure of the processed mRNA are added. As a result, paired
di-tags contain sequences from both termini of the same mRNA. Thus, the exact
transcription boundaries can be identified.
In combination with second generation sequencing methods, a vast increase in
throughput can be obtained (MS-PET) (Ng et al. 2006). The PETs have an average
length of 40 bp per di-tag and therefore two di-tags correspond to one sequencing
read of a first generation 454 sequencer.
It should be mentioned however, that SAGE and all derived methods involve
highly complex laboratory protocols. The increased throughput of next-generation
sequencing methods has shifted the bottleneck of sequencing to the adaptation of
these protocols. Also, it must be noted that the protocols are not all applicable to
prokaryotes, as the they are based on mRNA structures (5 capping, poly-A tail)
which are found in eukaryotes but not prokaryotes.
Shotgun-transcriptomics approaches are also suitable for the analysis of trancription. In principle, the approach is similar to the EST approach described before.
Messenger RNA is reverse transcribed into cDNA which is directly fragmented
and sequenced. Instead of using Sanger sequencing, high-throughput methods can
V. Mittard-Runte et al.
long chains of tags. In this classical protocol, the tags are then cloned into a plasmid vector and their sequence is obtained using Sanger sequencing. The specific
cleavage site remains in the DNA-vector and can be used as a separator.
If genomic sequences of the organism are known, either in the form of ESTs or
as whole genome, the tags obtained from sequencing can be mapped onto the longer
sequences, thereby enabling a quantitative analysis of differential gene-expression.
The number of similar tags found can be assumed to correlate (although not linearly
due to the involvement of PCR-amplification) with the original amount of cDNA
fragments from which these tags are uniquely derived.
If no sequence information is available for the source organism, the sequence
tags can still be analysed quantitatively, as the tag sequence is an (almost unique)
identifier for the corresponding mRNA. However, mapping of short tags of 11 bp to
available databases for sequence comparison will yield a large proportion of false
positive hits.
An initial problem of the SAGE protocol was the high amount of source poly-A
RNA required (approximately 2.5–5 μg). Therefore, Datson et al. (1999) developed
MicroSAGE which uses strepatividin coated tubes instead of beads. The method
requires only 1 ng of RNA, which is approximately the equivalent of 100,000
mammalian cells
Another drawback of the initial SAGE approach was the short tag-length. Several
enhancements of the technique have been published which provide an improved taglength by using different restriction enzymes. Tag lengths of 20 bp (LongSAGE)
(Saha et al. 2002), and 26 bp (SuperSAGE) (Matsumura et al. 2003), have been
reported. These provide increasingly significant database hits.
Another improvement has been the development of gene identification methods
based on paired-end di-tags (GIS-PETs) (Ng et al. 2005). In contrast to SAGE, PETs
are derived from the 3 and 5 signatures of mRNA. For the 3 end, primers specific to the capping structure of the processed mRNA are added. As a result, paired
di-tags contain sequences from both termini of the same mRNA. Thus, the exact
transcription boundaries can be identified.
In combination with second generation sequencing methods, a vast increase in
throughput can be obtained (MS-PET) (Ng et al. 2006). The PETs have an average
length of 40 bp per di-tag and therefore two di-tags correspond to one sequencing
read of a first generation 454 sequencer.
It should be mentioned however, that SAGE and all derived methods involve
highly complex laboratory protocols. The increased throughput of next-generation
sequencing methods has shifted the bottleneck of sequencing to the adaptation of
these protocols. Also, it must be noted that the protocols are not all applicable to
prokaryotes, as the they are based on mRNA structures (5 capping, poly-A tail)
which are found in eukaryotes but not prokaryotes.
Shotgun-transcriptomics approaches are also suitable for the analysis of trancription. In principle, the approach is similar to the EST approach described before.
Messenger RNA is reverse transcribed into cDNA which is directly fragmented
and sequenced. Instead of using Sanger sequencing, high-throughput methods can
