transcribed and translated do not produce detectable transcripts or proteins. Conversely, some sequences that do not have “gene” features are expressed. It would be
easy to say that sequences should only be considered genes if there is protein
evidence showing that they are translated, but “comprehensive” proteomics data
are difficult (perhaps impossible?) and expensive to obtain, and thus many genomes
do not have substantive proteomics support. Transcriptome data can also be used to
identify which genes are transcribed, but not all transcripts are translated. Additionally, some genes may only be active under very rare circumstances which might
make discovery of their corresponding transcripts and/or proteins nearly impossible.
Consequently, genes are typically identified using homology-based and ab initio
approaches, with most/many reported genes actually being “gene models.” Lastly,
structural and regulatory RNA molecules perform functions analogous to structural
proteins and transcription regulators, so perhaps they should be considered genes.
For the “high-quality” plant genomes shown in Table 3, reported gene numbers
range from 7,640 to 114,227, with an average gene number of 37,980.
14 Gymnosperm genome gene estimates, which are not in Table 3 based on the relatively low
quality of their genome assemblies, are 16,386 for Picea glauca, 27,491 for Gnetum
montanum, 50,172 for Pinus taeda, and 70,968 for Picea abies (Warren et al. 2015a;
Neale et al. 2014; Wan et al. 2018; Nystedt et al. 2013). These values fit comfortably
within the gene number range from Table 3. Of note, the chlorophyte algae have the
smallest gene numbers. While the smaller numbers may reflect one of the issues
mentioned below, it is likely that the numbers are at least partly reflective of the
tremendous amount of evolutionary time separating chlorophytes from the
angiosperms.
So what accounts for this variation in gene numbers among the seed plants?
Undoubtedly polyploidy and segmental duplication accounts for some of the differences. However, the seemingly wide range in gene counts is most likely due to the
fact that identification of genes using computer algorithms is highly dependent upon
the quality of the algorithms, the parameters used in gene searching, and the quality
and availability of previously annotated “training” sequences (Michael and Jackson
2013). Two highly competent research groups can come up with very different
results even with only seemingly modest differences in their approaches. For
example, in an analysis of 100 BAC clones from loblolly pine, Wegrzyn et al.
(2013) detected only one gene in the BAC clone set using MAKER (Holt and
Yandell 2011) informed by Augustus (Stanke et al. 2004) and the eukaryotic
orthologous proteins (KOGS) database (Koonin et al. 2004). However, our analysis
of the same group of BACs using MAKER informed by Augustus, GeneMark (Isono
et al. 1994), FGENESH (Solovyev et al. 2006), and cDNA data resulted in identification of 154 gene models of which 10 had enough ancillary scientific support to be
declared “certified genes” (Perera et al. 2018). Clearly, which combinations of
algorithms and databases are utilized is not trivial. Moreover, what gene number is
14 In their review of the 50 first plant genomes, Michael and Jackson (2013) report a range in gene
numbers between 20,000 and 94,000 with an average of 32,605.
Sequencing Plant Genomes
171
Précédent

- 180/342

Suivant