assemblies. As of October 10, 2017, there are 72 plant reference-quality genome
assemblies in NCBI.
6.2 The Myth of Plant Genome Sequences
On more or less a biweekly basis, another plant genome paper is published. Many of
these papers have the phrase “whole genome sequence” in their titles which has a
somewhat misleading effect of suggesting that the entire genome of each of these
plants is complete. How many plant genomes have been completely sequenced and
assembled without mistakes or holes? The answer, like the majority of biological
questions, is probably unknowable. With this said, there are, as of this writing, two
plant genomes that appear to be finished – specifically, the chlorophyte green algae
species Ostreococcus lucimarinus and Micromonas commode (Palenik et al. 2007;
Worden et al. 2009). The tiny genomes of these algae (13.2 and 29 Mb, respectively)
and their relatively low levels of repeat sequences allowed production of ostensibly
complete genome assemblies. However, all other plant genomes currently listed in
NCBI Genome possess gaps between scaffolds and regions that have been recalcitrant to chromosome/linkage group assignment (Table 3). These gaps are due to
problems associated with assembling segmental duplications and highly similar
repeat sequences. Moreover, centromeres are missing from all land plant genomes,
with the partial exception of rice (Yan and Jiang 2007), due to their high quantities of
repeats. This is unfortunate as some centromeric regions have been shown to contain
important genes (Wu et al. 2011; Fan et al. 2011; Shahinnia et al. 2012; Chen et al.
2013b; Yan and Jiang 2007). Regardless, even the best land plant genome sequences
contain gaps; as mentioned above, even the Arabidopsis thaliana genome appears to
be only 76% complete.
6.3 The Un- and Underrepresented
Examination of the species listed in Table 3 reveals that all the plant species with
reference-quality genome sequences are angiosperms or chlorophyte algae. The
angiosperms are the most species-rich plant clade and are the source of most
human calories. Consequently, the high representation of angiosperms (especially
those species of agronomic value) in Table 3 is not surprising. The chlorophyte algae
have relatively small genomes and have been touted as potential tools in bioenergy
production which explains their representation in the table. Some of the groups that
do not have a reference-quality genome representative include:
• Charophyte algae – Land plants (embryophytes) emerged as a subgroup of the
Charophyte algae. Thus Charophyte genomes represent a logical outgroup for
exploring the early evolution of land plants (de Vries and Archibald 2018). To
166
D. G. Peterson and M. Arick
assemblies in NCBI.
6.2 The Myth of Plant Genome Sequences
On more or less a biweekly basis, another plant genome paper is published. Many of
these papers have the phrase “whole genome sequence” in their titles which has a
somewhat misleading effect of suggesting that the entire genome of each of these
plants is complete. How many plant genomes have been completely sequenced and
assembled without mistakes or holes? The answer, like the majority of biological
questions, is probably unknowable. With this said, there are, as of this writing, two
plant genomes that appear to be finished – specifically, the chlorophyte green algae
species Ostreococcus lucimarinus and Micromonas commode (Palenik et al. 2007;
Worden et al. 2009). The tiny genomes of these algae (13.2 and 29 Mb, respectively)
and their relatively low levels of repeat sequences allowed production of ostensibly
complete genome assemblies. However, all other plant genomes currently listed in
NCBI Genome possess gaps between scaffolds and regions that have been recalcitrant to chromosome/linkage group assignment (Table 3). These gaps are due to
problems associated with assembling segmental duplications and highly similar
repeat sequences. Moreover, centromeres are missing from all land plant genomes,
with the partial exception of rice (Yan and Jiang 2007), due to their high quantities of
repeats. This is unfortunate as some centromeric regions have been shown to contain
important genes (Wu et al. 2011; Fan et al. 2011; Shahinnia et al. 2012; Chen et al.
2013b; Yan and Jiang 2007). Regardless, even the best land plant genome sequences
contain gaps; as mentioned above, even the Arabidopsis thaliana genome appears to
be only 76% complete.
6.3 The Un- and Underrepresented
Examination of the species listed in Table 3 reveals that all the plant species with
reference-quality genome sequences are angiosperms or chlorophyte algae. The
angiosperms are the most species-rich plant clade and are the source of most
human calories. Consequently, the high representation of angiosperms (especially
those species of agronomic value) in Table 3 is not surprising. The chlorophyte algae
have relatively small genomes and have been touted as potential tools in bioenergy
production which explains their representation in the table. Some of the groups that
do not have a reference-quality genome representative include:
• Charophyte algae – Land plants (embryophytes) emerged as a subgroup of the
Charophyte algae. Thus Charophyte genomes represent a logical outgroup for
exploring the early evolution of land plants (de Vries and Archibald 2018). To
166
D. G. Peterson and M. Arick
