carry megabase-sized DNA inserts, roughly 50% of YAC inserts were chimeric
or were characterized by insert rearrangements. In contrast, BAC inserts were
highly stable. Moreover, isolating YACs was relatively difficult, while BACs
were as easy to isolate as any single-copy bacterial plasmid (see Peterson et al.
2000 for review). Additionally, BACs could be readily end-sequenced, and
BAC-end sequences were a highly efficient way to span gaps (up to
100–300 kb) during genome assembly.
• The design and production of automated Sanger sequencers by Lloyd Smith,
Leroy Hood, and Applied Biosystems (Smith et al. 1986; Shendure et al. 2017) –
Automation greatly reduced the human hours involved in actual sequencing and
post-sequencing data processing.
• The inclusion of random fragment (“shotgun”) sequencing in the genome
sequencing process – previously considered a complete waste of money, production of shotgun sequence reads became a feasible approach for helping improve
assembly of non-repetitive regions of the genome due to the drop in sequencing
costs.
• Algorithm development and construction of computational tools to expedite
many aspects of the sequencing, assembly, and annotation process, notably
Phred, Phrap, Consed, the TIGR assembler, and the Celera assembler (Shendure
et al. 2017) – The bioinformatics tools utilized by the HGP were adopted by the
plant genome community.
5.2 BAC-by-BAC Sequencing
The first plant to have its genome sequenced was Arabidopsis thaliana, an angiosperm dicot from the Brassicaceae. This mustard plant was selected not for its
socioeconomic importance but because it was an established model for plant physiology and genetics and had all the makings of a superlative genomics model. In
brief, it (a) has a small genome size, (b) possesses a rapid life cycle, (c) is short and
compact and needs little space to grow, (d) can be crossed and selfed, (e) produces
numerous seeds, (f) can be transformed, and (g) can be readily used to create mutants
using a number of methods, with T-DNA mutagenesis being particularly valuable
(see Koornneef and Meinke 2010 for review).
The second plant genome project focused on rice (Oryza sativa spp. japonica), an
angiosperm monocot with C3 carbon fixation. Rice has many features of a model
plant (small genome size), but it also has the added advantage of being the single
most important source of human calories (Awika 2011).
Arabidopsis and rice were sequenced and assembled using a BAC-by-BAC
strategy (Claros et al. 2012; Jackson 2016) that paralleled the successful strategy
employed by the publicly funded HGP. The tenets of the BAC-by-BAC strategy are
as follows, although neither the Arabidopsis nor rice genomes were sequenced
exactly as described below:
152
D. G. Peterson and M. Arick
Précédent

- 161/342

Suivant