74
Box 1: Nucleic Acid Sequence Analysis
Background
The nucleic acids contain information in the shape
of a code constituting of two purines, Adenine A and
Guanine G, and two pyrimidine bases, Cytosine C and
either Thymine T in DNA or Uracil U in RNA. Selective
pairing of A with T and G with C gives rise to the stable double strand structure of DNA and confers a
mechanism to pass on the information in the coded
sequence via polymerization, i.e. DNA replication and
RNA transcription (in this case, substituting U for T)
(Alberts et al. 2008; Klug et al. 2012). The widely
applied DNA/RNA sequencing methods to read the
nucleotide code are based on this selective binding.
The first sequencing developed by Sanger in the
1970s required four separate polymerizations, each
with a fraction of dideoxynucleotides (ddNTPs) which
would terminate the elongation – hence the name
‘chain-termination method’ (Lu et al. 2016). Parallel
size separation (using gel-electrophoresis) of the synthesized strands, each with a specific dd-nucleotide at
the end, and subsequent radioactive detection allowed
to infer the order of the different bases in the template’s sequence. Modern techniques for Sanger
sequencing are based on fluorescently labeled
ddNTPs, emitting differentiable signals, which can be
detected by a laser and evaluated electronically
(Schuster 2008). Recently developed second-generation sequencing (such as Illumina) use dNTPs which
emit a base-specific fluorescent signal when the phosphordiester bond is formed and the DNA elongated.
Different to the traditional Sanger sequencing, the
process does not require termination and every elongation process yields a signal per nucleic acid. The
advantages of these sequencing methods lie in highthroughput through the simultaneous sequencing of
multiple DNA/RNA fragments (e.g., from environmental samples) from a variety of organisms with usually reliable high-quality results (Schuster 2008). The
drawbacks belay in comparatively short sequence
strands (about 100–300 bp), demanding assemblies to
solve the ‘puzzle’ of different short fragments.
However, third-generation sequencing (such as
offered by PacBio with the SMRT cell) make use of
double-stranded DNA with two hairpin structures at
the end, the so-called SMRTbell. This way, fragments
of several thousand base pairs may be sequenced,
which may subsequently be complemented by shorter
fragments to maintain the quality standard via high
coverage (Rhoads and Au 2015).
The emerging fourth generation sequencing technique, the nanopore sequencing (such as the MiniION
by Oxford Nanopore Techniques), does not require
previous amplification but aims at directly sequencing
single molecules and promises to sequence tens of
kilobases (kb). A membrane is equipped with nanopores that is selectively permeable for DNA and
RNA. An electric force is driving the electrophoresis
of the negatively charged fragments towards the anode
and, thus, into the membrane. A motor protein is ratcheting the fragment through the membrane. This causes
different perturbations of the membrane current
depending on the nucleotide, which may be computationally translated into base sequences (Cherf et al.
2012; Feng et al. 2015). Different from previous
sequencing methods, the fourth generation nanopore
sequencing may even be used to analyze proteins,
polymers, and other single-strand macromolecules
(Feng et al. 2015).
Strategies
To target a particular portion of the queried nucleotide sequence, e.g., targeting the 16S rRNA/rDNA of
microorganisms for phylogenetic assessment, specific
primer sequences can be used.
A variety of techniques grouped under the description of restriction site-associated DNA sequencing
(RADseq) is currently in scope for assessing genotypic
differences of a range of organisms, including those
with largely unknown genomes. These techniques are
based on digestion of isolated DNA with one or few
restriction enzymes and subsequent sequencing of
resulting fragments. As most restriction sites prevail
among specimen and closely related species, predominantly similar sets of loci are sequenced, at which different alleles can be identified (Andrews et al. 2016).
In case of whole genome sequencing using NGS,
short fragments of DNA of few hundred base pair
length are inserted into vectors, called library. To aid in
later assembly, libraries with shotgun mate pair fragments of specified greater lengths complement the
short vector sequences, which consist of a high fragment coverage. After standard quality controls of the
reads (including adapter and primer removal), the
assembly of the genome from the multitude of small
sequences relies on overlapping regions and mate pairs
(e.g., Baumgarten et al. 2015).
Prior to RNA sequencing, the RNA-template has to
be transcribed into a cDNA, using a reverse transcriptase. A quantitative interpretation of transcriptome and
(continued)
J. D. Brüwer and H. Buck-Wiese
Box 1: Nucleic Acid Sequence Analysis
Background
The nucleic acids contain information in the shape
of a code constituting of two purines, Adenine A and
Guanine G, and two pyrimidine bases, Cytosine C and
either Thymine T in DNA or Uracil U in RNA. Selective
pairing of A with T and G with C gives rise to the stable double strand structure of DNA and confers a
mechanism to pass on the information in the coded
sequence via polymerization, i.e. DNA replication and
RNA transcription (in this case, substituting U for T)
(Alberts et al. 2008; Klug et al. 2012). The widely
applied DNA/RNA sequencing methods to read the
nucleotide code are based on this selective binding.
The first sequencing developed by Sanger in the
1970s required four separate polymerizations, each
with a fraction of dideoxynucleotides (ddNTPs) which
would terminate the elongation – hence the name
‘chain-termination method’ (Lu et al. 2016). Parallel
size separation (using gel-electrophoresis) of the synthesized strands, each with a specific dd-nucleotide at
the end, and subsequent radioactive detection allowed
to infer the order of the different bases in the template’s sequence. Modern techniques for Sanger
sequencing are based on fluorescently labeled
ddNTPs, emitting differentiable signals, which can be
detected by a laser and evaluated electronically
(Schuster 2008). Recently developed second-generation sequencing (such as Illumina) use dNTPs which
emit a base-specific fluorescent signal when the phosphordiester bond is formed and the DNA elongated.
Different to the traditional Sanger sequencing, the
process does not require termination and every elongation process yields a signal per nucleic acid. The
advantages of these sequencing methods lie in highthroughput through the simultaneous sequencing of
multiple DNA/RNA fragments (e.g., from environmental samples) from a variety of organisms with usually reliable high-quality results (Schuster 2008). The
drawbacks belay in comparatively short sequence
strands (about 100–300 bp), demanding assemblies to
solve the ‘puzzle’ of different short fragments.
However, third-generation sequencing (such as
offered by PacBio with the SMRT cell) make use of
double-stranded DNA with two hairpin structures at
the end, the so-called SMRTbell. This way, fragments
of several thousand base pairs may be sequenced,
which may subsequently be complemented by shorter
fragments to maintain the quality standard via high
coverage (Rhoads and Au 2015).
The emerging fourth generation sequencing technique, the nanopore sequencing (such as the MiniION
by Oxford Nanopore Techniques), does not require
previous amplification but aims at directly sequencing
single molecules and promises to sequence tens of
kilobases (kb). A membrane is equipped with nanopores that is selectively permeable for DNA and
RNA. An electric force is driving the electrophoresis
of the negatively charged fragments towards the anode
and, thus, into the membrane. A motor protein is ratcheting the fragment through the membrane. This causes
different perturbations of the membrane current
depending on the nucleotide, which may be computationally translated into base sequences (Cherf et al.
2012; Feng et al. 2015). Different from previous
sequencing methods, the fourth generation nanopore
sequencing may even be used to analyze proteins,
polymers, and other single-strand macromolecules
(Feng et al. 2015).
Strategies
To target a particular portion of the queried nucleotide sequence, e.g., targeting the 16S rRNA/rDNA of
microorganisms for phylogenetic assessment, specific
primer sequences can be used.
A variety of techniques grouped under the description of restriction site-associated DNA sequencing
(RADseq) is currently in scope for assessing genotypic
differences of a range of organisms, including those
with largely unknown genomes. These techniques are
based on digestion of isolated DNA with one or few
restriction enzymes and subsequent sequencing of
resulting fragments. As most restriction sites prevail
among specimen and closely related species, predominantly similar sets of loci are sequenced, at which different alleles can be identified (Andrews et al. 2016).
In case of whole genome sequencing using NGS,
short fragments of DNA of few hundred base pair
length are inserted into vectors, called library. To aid in
later assembly, libraries with shotgun mate pair fragments of specified greater lengths complement the
short vector sequences, which consist of a high fragment coverage. After standard quality controls of the
reads (including adapter and primer removal), the
assembly of the genome from the multitude of small
sequences relies on overlapping regions and mate pairs
(e.g., Baumgarten et al. 2015).
Prior to RNA sequencing, the RNA-template has to
be transcribed into a cDNA, using a reverse transcriptase. A quantitative interpretation of transcriptome and
(continued)
J. D. Brüwer and H. Buck-Wiese
