computing environments; however, there is still much room for improvement.
Whether across campus or in the cloud, HPC/parallel processing is necessary for
timely analysis of genomic data (see Ocana and de Oliveira 2015 for review).
A discussion of all the bioinformatics tools, which can be utilized in DNA
sequence assembly, annotation, comparison, etc., is beyond the scope of this chapter.
However, Table 2 presents a brief description, references, and download links to the
major sequence analysis tools currently being utilized by the research team at the
Institute for Genomics, Biocomputing & Biotechnology (IGBB) at Mississippi State
University.
With regard to the use of bioinformatics tools, we do have a few suggestions
based on our experiences. These suggestions are:
• Be wary when using commercial software. Some commercial software is good,
but much of it is not. In our experience, many commercial bioinformatics programs aren’t capable of handling large volumes of NGS data. One commercial
program we tested, which was run using the desktop computer provided with the
software, took several months to produce a bacterial DNA assembly; we assembled the same sequence reads using the publicly available ABySS (Jackman et al.
2017) and Velvet (Zerbino and Birney 2008) programs independently on an HPC
system. Both the ABySS and Velvet analyses took less than 5 min. The assembly
produced by the commercial software was aesthetically pleasing (there were no
gaps), but the assembly looked nothing like the highly similar ABySS and Velvet
assemblies. Moreover, the commercial software produced an assembly that shared
no synteny with published genomes of species from the same bacterial genus.
Because of its proprietary nature, we could not figure out why the commercial
program did what it did. In short, if you can’t look behind the curtains to see what
is going on, you better have good reason to trust the commercial software vendor.
• Use Linux and its default command line (terminal) interface. Most bioinformatics
tools are written to operate in a Linux environment, and very few have a GUI
(graphic user interface) available. Most HPC centers have Linux operating
systems.
• Adopt the literate programming model. Literate programming (Knuth 1984) is
the practice of placing documentation and source code in the same document. The
value of this is that a single document has the what, why, and how information on
a particular script. Adding a summary of results to the document increases its
usefulness further. Org-mode (https://orgmode.org), an Emacs text processor
extension, is an excellent tool for this type of documentation.
• Document and track everything. Experimenting with different tools and parameters is vital to produce the best possible analysis from the data. Even changing
the versions of tools can have a significant impact on the results (Salzberg et al.
2012; Conesa et al. 2016). With all this variability, documentation is crucial to the
reproducibility of any analysis. To aid in this task, a version control system can be
utilized to track the changes through time. Git (https://git-scm.com/) is a popular
open-source distributed version control system and the one we use in our lab.
148
D. G. Peterson and M. Arick
Précédent

- 157/342

Suivant