DNA must be genes. We now know that differences in genome size typically are
reflective of differences in repetitive sequence contents and, to a much lesser extent,
polyploidy and gene duplication (see Peterson 2014, for review). The reasons why
natural selection has resulted in obese and skinny genomes is not well understood,
but the nature of the differences in genome sizes has been elucidated. Consequently,
it has been recommended that the C-value paradox should now be referred to as the
C-value enigma (Gregory 2005).
2.6 Homogeneity
When generating a genome sequence for a species or clade, it might seem logical to
include DNA from many different individuals as, perhaps, a means of looking at
diversity. This is absolutely incorrect. The assembly of sequence reads into long
DNA stretches is most easily performed if there is a complete absence of heterogeneity at any locus (or if this is not an option, as little heterogeneity as possible).
Ideally a reference genome is prepared from a source of identical haploid cells (e.g.,
a megagametophyte), a doubled-haploid cell line, or a highly inbred individual. The
goal of producing a genome sequence is to produce the most complete representation
of a single (haploid) genome as possible. The sequence itself provides no information on genetic diversity among individuals, but once completed, the genome serves
as a way in which diversity can be quickly assessed (see Sect. 3.18 below). One way
to circumvent problems with heterozygosity when DNA must be obtained from a
diploid, heterozygous individual is through a process known as phasing (see Sect.
3.23 below).
2.7 Chromatin
The mixture of DNA and its associated proteins is called chromatin. In eukaryotic
cells, DNA is always found within the context of chromatin and chromosomes
(Fig. 4). On a particular DNA molecule, two loci may be separated by thousands
or millions of base pairs, but due to chromatin looping/folding, these loci may be
very near to each other in the 3D environment of the nucleus (Fig. 4f). Since
genomes evolved within the context of chromatin, it is perhaps not surprising that
regulatory DNA sequences may be close to the gene they regulate in 3D space while
being distant in base pairs from that gene.
Chromatin from each chromosome occupies a relative distinct domain within the
interphase nucleus. More recently, chromatin conformation capture techniques have
been coupled with NGS to look at pairwise associations of sequences in nuclei (see
Sect. 3.24 below for more details). This research has led to the discovery of
topologically associating domains (TADs) which are collections of DNA sequences
that tend to be closely associated (and ostensibly in contact) with each other in the
Sequencing Plant Genomes
119
reflective of differences in repetitive sequence contents and, to a much lesser extent,
polyploidy and gene duplication (see Peterson 2014, for review). The reasons why
natural selection has resulted in obese and skinny genomes is not well understood,
but the nature of the differences in genome sizes has been elucidated. Consequently,
it has been recommended that the C-value paradox should now be referred to as the
C-value enigma (Gregory 2005).
2.6 Homogeneity
When generating a genome sequence for a species or clade, it might seem logical to
include DNA from many different individuals as, perhaps, a means of looking at
diversity. This is absolutely incorrect. The assembly of sequence reads into long
DNA stretches is most easily performed if there is a complete absence of heterogeneity at any locus (or if this is not an option, as little heterogeneity as possible).
Ideally a reference genome is prepared from a source of identical haploid cells (e.g.,
a megagametophyte), a doubled-haploid cell line, or a highly inbred individual. The
goal of producing a genome sequence is to produce the most complete representation
of a single (haploid) genome as possible. The sequence itself provides no information on genetic diversity among individuals, but once completed, the genome serves
as a way in which diversity can be quickly assessed (see Sect. 3.18 below). One way
to circumvent problems with heterozygosity when DNA must be obtained from a
diploid, heterozygous individual is through a process known as phasing (see Sect.
3.23 below).
2.7 Chromatin
The mixture of DNA and its associated proteins is called chromatin. In eukaryotic
cells, DNA is always found within the context of chromatin and chromosomes
(Fig. 4). On a particular DNA molecule, two loci may be separated by thousands
or millions of base pairs, but due to chromatin looping/folding, these loci may be
very near to each other in the 3D environment of the nucleus (Fig. 4f). Since
genomes evolved within the context of chromatin, it is perhaps not surprising that
regulatory DNA sequences may be close to the gene they regulate in 3D space while
being distant in base pairs from that gene.
Chromatin from each chromosome occupies a relative distinct domain within the
interphase nucleus. More recently, chromatin conformation capture techniques have
been coupled with NGS to look at pairwise associations of sequences in nuclei (see
Sect. 3.24 below for more details). This research has led to the discovery of
topologically associating domains (TADs) which are collections of DNA sequences
that tend to be closely associated (and ostensibly in contact) with each other in the
Sequencing Plant Genomes
119
