4 Bioinformatics and Computer Concerns
In the early days of genome analysis, there were relatively few computer scientists
working on issues related to DNA sequencing, and consequently, many of the first
sequence analysis scripts were written by biologists. While there are likely some
major exceptions, biologists, in general, do not write computer code as efficiently
and effectively as those trained specifically to do so. Co-author Peterson once was
involved in a discussion with J. Craig Venter, a key (and controversial) figure in the
Human Genome Project (HGP). When Dr. Venter was asked what was the biggest
scientific breakthrough that allowed the HGP to finish under budget and ahead of
schedule (Hilgartner 2017), he said without hesitation, “Hiring computer programmers to rewrite our scripts.”
The HGP led to considerable improvement in sequence manipulation/analysis
tools. The HGP and the plant and animal genome projects that followed were
dependent upon on the use of such tools including the Phred-Phrap-Consed suite
(Nickerson et al. 1997), BLAST (Altschul et al. 1990) and its derivatives,
RepeatMasker (Smit et al. 2013), and a number of gene/repeat finding scripts.
Computer clusters played a large role in sequence manipulation and analysis,
although much work was done on individual workstations.
When NGS approaches emerged on the scene, they resulted in such a glut of data
that many of the sequencing centers that were beta-testing the machines quickly
filled up their computer storage capacities. These researchers were accustomed to
saving the electropherogram (raw) data produced by Sanger sequencers in case
reanalysis was determined prudent, and they initially tried saving the massive
image files produced by 454 and Illumina machines in the same manner. There
was not room for these large data files to be stored or archived in NCBI/DDBJ/
EMBL as was common practice with Sanger data. Consequently, throwing away raw
data files, which was considered very bad form in the pre-NGS days, became a
necessity (Richter and Sexton 2009).
NGS has forced genome scientists to start utilizing the power of highperformance computing (HPC). This is an ongoing transition that has not been
pain-free. In brief, few early bioinformatics programs were developed for HPC
environments. While the programs could often be run on HPC systems, they were
not designed to take advantage of the efficiency of HPC systems. Thus running early
bioinformatics scripts on a supercomputer was not a huge step up from running
the programs on the desktop workstations on which they often were designed. The
problem of script inefficiency was partly offset by creation of pipelines where the
results of an analysis by one program were automatically fed (by customized scripts)
to the next program in a chain of programs. Pipelines are useful even when scripts are
written for HPC. However, a major problem in bioinformatics has been a failure to
write/adapt scripts to utilize the many advantages of parallel computing, the division
of computational labor between multiple (often hundreds or thousands of) processors
rather than relying on one high RAM machine or small cluster for analysis. Today,
more bioinformaticians are writing new and adapting old tools to utilize parallel
Sequencing Plant Genomes
147
Précédent

- 156/342

Suivant