sequenced from 5
0 to 3
0 , the two primers used were necessarily complementary to
different strands of the vector molecule. If the insert size was short enough (e.g.,
1,400 bp), the forward and reverse reads might overlap, allowing generation of
consensus sequence. If the insert was larger, the paired-end sequences would be
separated by a gap of unknown sequence. If the gap size is known, the paired-end
sequences can be used as bridges between contigs. For example, let’s say you have a
BAC with an insert size of 200 kb and you have sequenced 700 nt from each end of
the insert. Both the resulting sequence data and positional data can be of considerable use. If alignment of the paired-end sequences (in this case, BAC-end sequences
or BESs) with contigs reveals that each BES is found in a different contig, those
contigs can be grouped into a scaffold with an estimated space of 200 kb between the
BES sites. To our knowledge, BAC-end sequencing is only performed using automated Sanger sequencing.
When NGS technologies were being developed, it was clear that being able to
produce paired-end sequences would be desirable, primarily as a means of bridging
gaps and increasing output. For Illumina, development of paired-end sequencing
was relatively straightforward. As shown in Fig. 5m, n, after bridge amplification,
reverse strands are cut from the slide leaving only forward strands to serve as
templates for sequencing. However, after sequencing using the forward strands,
the 3
0 ends of the forward strands can be unblocked, and bridge amplification can
be used to regenerate reverse strands. The forward strands can then be cleaved from
the slide, the 3
0 ends of the reverse strands can be blocked, and sequencing using the
reverse strands as templates can be conducted. However, target molecules longer
200 + 100 + 60 + 40 + 30 + 30 + 20 + 20 + 20 + 20 + 18 + 9 + 9 + 7 + 7 + 7 + 7 + 5 + 5 + 2 + 2 + 2 = 620 kb
N50 = 60 kb
a
b
c
d
e
40
20
18
30 30
20 20 20
200 kb
100
60
40
20
18
30 30
20 20 20
200 kb
100
60
b
k
0
1
3
b
k
0
1
3
Fig. 8 Understanding N50. (a) Unordered contigs representing a genome assembly. Each contig is
represented by a rectangle with its width being proportional to the contig’s relative length in kb. (b)
In silico, the contigs are arranged from longest to shortest. (c) The sum length of the contigs is
determined (620 kb), and (d) 50% of the sum contig length is calculated (310 kb). (d) Starting from
either end of the ordered contigs and moving 50% of the sum contig length (i.e., 310 kb) will bring
you to a particular contig (arrow between braces). (e) The length of the contig intersected by the
arrow (i.e., 60 kb) is the N50. In other words, 50% of the sum contig length is contained in contigs
!60 kb in length. Note that N50 values can be reported for scaffolds as well as contigs
Sequencing Plant Genomes
141
0 to 3
0 , the two primers used were necessarily complementary to
different strands of the vector molecule. If the insert size was short enough (e.g.,
1,400 bp), the forward and reverse reads might overlap, allowing generation of
consensus sequence. If the insert was larger, the paired-end sequences would be
separated by a gap of unknown sequence. If the gap size is known, the paired-end
sequences can be used as bridges between contigs. For example, let’s say you have a
BAC with an insert size of 200 kb and you have sequenced 700 nt from each end of
the insert. Both the resulting sequence data and positional data can be of considerable use. If alignment of the paired-end sequences (in this case, BAC-end sequences
or BESs) with contigs reveals that each BES is found in a different contig, those
contigs can be grouped into a scaffold with an estimated space of 200 kb between the
BES sites. To our knowledge, BAC-end sequencing is only performed using automated Sanger sequencing.
When NGS technologies were being developed, it was clear that being able to
produce paired-end sequences would be desirable, primarily as a means of bridging
gaps and increasing output. For Illumina, development of paired-end sequencing
was relatively straightforward. As shown in Fig. 5m, n, after bridge amplification,
reverse strands are cut from the slide leaving only forward strands to serve as
templates for sequencing. However, after sequencing using the forward strands,
the 3
0 ends of the forward strands can be unblocked, and bridge amplification can
be used to regenerate reverse strands. The forward strands can then be cleaved from
the slide, the 3
0 ends of the reverse strands can be blocked, and sequencing using the
reverse strands as templates can be conducted. However, target molecules longer
200 + 100 + 60 + 40 + 30 + 30 + 20 + 20 + 20 + 20 + 18 + 9 + 9 + 7 + 7 + 7 + 7 + 5 + 5 + 2 + 2 + 2 = 620 kb
N50 = 60 kb
a
b
c
d
e
40
20
18
30 30
20 20 20
200 kb
100
60
40
20
18
30 30
20 20 20
200 kb
100
60
b
k
0
1
3
b
k
0
1
3
Fig. 8 Understanding N50. (a) Unordered contigs representing a genome assembly. Each contig is
represented by a rectangle with its width being proportional to the contig’s relative length in kb. (b)
In silico, the contigs are arranged from longest to shortest. (c) The sum length of the contigs is
determined (620 kb), and (d) 50% of the sum contig length is calculated (310 kb). (d) Starting from
either end of the ordered contigs and moving 50% of the sum contig length (i.e., 310 kb) will bring
you to a particular contig (arrow between braces). (e) The length of the contig intersected by the
arrow (i.e., 60 kb) is the N50. In other words, 50% of the sum contig length is contained in contigs
!60 kb in length. Note that N50 values can be reported for scaffolds as well as contigs
Sequencing Plant Genomes
141
