data, and using Hi-C chromatin conformation capture and/or optical mapping
techniques to increase contig/scaffold accuracy. Because improvements in sequencing have occurred so quickly, bioinformatics tools for assembling, analyzing, and
annotating genomes have not been standardized; indeed, such tools can vary widely
with regard to input needs, output quality, and scalability. With hundreds of plant
genomes in various states of completion, much is being learned about general trends
in plant genome evolution – for example, the predominance of polyploidy and
paleopolyploidy events in angiosperms and the enormous contribution of LTR
retrotransposons to genome size expansions/contractions in land plants. Plant
genome sequencing projects have also made apparent the difficulties associated
with identifying and quantifying gene numbers. While genome sequences are
powerful tools, they are simply maps; they contain information about genomic
landmarks and their possible functions, but they are not equivalent to the complex
organisms they represent. Nonetheless, like any maps, they represent a means by
which one can explore more expeditiously. Integration of genome sequence maps
with chromatin configuration, biochemical, and developmental data is critical in
advancing understanding of how genes and genomes function and evolve in vivo.
1 Introduction
In each semester at the beginning of the first lecture of the graduate course “Genomes
and Genomics,” co-author Peterson asks his students the same question: “Why do
we do so much DNA sequencing?” Typically, the students look down at the floor
and sink lower in their chairs; they have sadly, and perhaps wisely, learned to avoid
trying to answer such open-ended questions, especially when it is clear that the
professor has a specific answer in mind. Eventually, he gives them the answer he
thinks is most accurate: “We sequence DNA because we can!”
Clearly researchers sequence genomes, transcriptomes, etc., for specific scientific
reasons, but generating nucleic acid sequences has become so fast and so inexpensive that scientists often use sequencing even when other approaches would afford
more direct answers. As shown in Fig. 1, the cost of sequencing DNA (machine time
and sequencing reagents only) has plummeted from $5,300 per megabase (Mb) in
2001 to $0.01 per Mb in 2016. For the cost of a case of pipet tips, a researcher can
generate more sequence data than they can manually annotate/examine in a lifetime.
However, even if a scientist is only interested in a small part of the sequence data
they generate (e.g., a gene family), by making the data available to the scientific
community through the US National Center for Biotechnology Information (NCBI),
the DNA Data Bank of Japan (DDBJ), or the European Molecular Biology
Laboratory-European Bioinformatics Institute (EMBL-EBI), the data producer is
Sequencing Plant Genomes
111
techniques to increase contig/scaffold accuracy. Because improvements in sequencing have occurred so quickly, bioinformatics tools for assembling, analyzing, and
annotating genomes have not been standardized; indeed, such tools can vary widely
with regard to input needs, output quality, and scalability. With hundreds of plant
genomes in various states of completion, much is being learned about general trends
in plant genome evolution – for example, the predominance of polyploidy and
paleopolyploidy events in angiosperms and the enormous contribution of LTR
retrotransposons to genome size expansions/contractions in land plants. Plant
genome sequencing projects have also made apparent the difficulties associated
with identifying and quantifying gene numbers. While genome sequences are
powerful tools, they are simply maps; they contain information about genomic
landmarks and their possible functions, but they are not equivalent to the complex
organisms they represent. Nonetheless, like any maps, they represent a means by
which one can explore more expeditiously. Integration of genome sequence maps
with chromatin configuration, biochemical, and developmental data is critical in
advancing understanding of how genes and genomes function and evolve in vivo.
1 Introduction
In each semester at the beginning of the first lecture of the graduate course “Genomes
and Genomics,” co-author Peterson asks his students the same question: “Why do
we do so much DNA sequencing?” Typically, the students look down at the floor
and sink lower in their chairs; they have sadly, and perhaps wisely, learned to avoid
trying to answer such open-ended questions, especially when it is clear that the
professor has a specific answer in mind. Eventually, he gives them the answer he
thinks is most accurate: “We sequence DNA because we can!”
Clearly researchers sequence genomes, transcriptomes, etc., for specific scientific
reasons, but generating nucleic acid sequences has become so fast and so inexpensive that scientists often use sequencing even when other approaches would afford
more direct answers. As shown in Fig. 1, the cost of sequencing DNA (machine time
and sequencing reagents only) has plummeted from $5,300 per megabase (Mb) in
2001 to $0.01 per Mb in 2016. For the cost of a case of pipet tips, a researcher can
generate more sequence data than they can manually annotate/examine in a lifetime.
However, even if a scientist is only interested in a small part of the sequence data
they generate (e.g., a gene family), by making the data available to the scientific
community through the US National Center for Biotechnology Information (NCBI),
the DNA Data Bank of Japan (DDBJ), or the European Molecular Biology
Laboratory-European Bioinformatics Institute (EMBL-EBI), the data producer is
Sequencing Plant Genomes
111
