Table 2 (continued)
Program name hyperlinked to program download site – description (reference)
GO Annotation Quality (GAQ) Score – calculates the quality of GO annotations associated with a
set of gene products (Buza et al. 2008)
a
GA2GEO – allows users to convert a file between the gene association format and an NCBI GEO
(Gene Expression Omnibus) type format
a
GOanna – allows users to quickly add more GO annotations by transferring GO annotations from
similar sequences (McCarthy et al. 2006)
a
GOanna2ga – used to convert the output from GOanna into a gene association file format
(McCarthy et al. 2011)
a
GOModeler – allows users to conduct a hypothesis-driven interrogation of their functional
genomics datasets (Manda et al. 2010)
a
GORetriever – returns GO annotations for a list of protein accessions or GO IDs. Used to retrieve
existing GO annotation (McCarthy et al. 2006)
a
HISAT2 – a fast and sensitive alignment program for mapping next-generation sequencing reads
(Kim et al. 2015)
HMMER – used for searching sequence databases for homologs of protein sequences and for
making protein sequence alignments. It implements methods using probabilistic models called
profile hidden Markov models (profile HMMs) (Finn et al. 2015)
InterProScan (iprscan) – scans sequences for matches against the InterPro collection of protein
signature databases (Zdobnov and Apweiler 2001)
LINKS – long interval nucleotide k-mer scaffolder (Warren et al. 2015b)
MAKER2 – a portable and easily configurable genome annotation pipeline (Cantarel et al. 2008;
Holt and Yandell 2011)
MGEScan-LTR – identifies long terminal repeats (LTR) (Rho et al. 2007)
Minimap2 – a versatile pairwise aligner for genomic and spliced nucleotide sequences (Li 2016)
Miniasm – ultrafast de novo assembly for long noisy reads (though having no consensus step)
(Li 2016)
MCScanX – algorithm for detection of synteny and collinearity (Wang et al. 2012b)
MUMmer – a system for rapidly aligning entire genomes, whether in complete or draft form
(Marcais et al. 2018)
MUSCLE – multiple sequence alignment (Edgar 2004)
NextClip – a tool for comprehensive quality analysis and read preparation for Nextera Long Mate
Pair (LMP) libraries (Leggett et al. 2014)
Nseg – used to mask nucleic acid sequences, needed by RepeatScout
OMSSA – open mass spectrometry search algorithm (Geer et al. 2004)
OrthoCluster – synteny block identification (Ng et al. 2009)
Picard – a set of tools (in Java) for working with next-generation sequencing data in the BAM
(http://www.htslib.org/) format
Pilon – an automated genome assembly improvement and variant detection tool (Walker et al.
2014)
ProQuant – tool for label-free protein quantification of SEQUEST output for the Microsoft
Windows environment (Bridges et al. 2007)
a
Proteogenomic mapping tool – provides experimentally based structural annotations at a complete
genome level (Sanders et al. 2011)
a
ProtIDer – enhances proteomics based on EST or EST assemblies by generating databases of
matching highly homologous proteins (McCarthy et al. 2006)
a
(continued)
150
D. G. Peterson and M. Arick
Program name hyperlinked to program download site – description (reference)
GO Annotation Quality (GAQ) Score – calculates the quality of GO annotations associated with a
set of gene products (Buza et al. 2008)
a
GA2GEO – allows users to convert a file between the gene association format and an NCBI GEO
(Gene Expression Omnibus) type format
a
GOanna – allows users to quickly add more GO annotations by transferring GO annotations from
similar sequences (McCarthy et al. 2006)
a
GOanna2ga – used to convert the output from GOanna into a gene association file format
(McCarthy et al. 2011)
a
GOModeler – allows users to conduct a hypothesis-driven interrogation of their functional
genomics datasets (Manda et al. 2010)
a
GORetriever – returns GO annotations for a list of protein accessions or GO IDs. Used to retrieve
existing GO annotation (McCarthy et al. 2006)
a
HISAT2 – a fast and sensitive alignment program for mapping next-generation sequencing reads
(Kim et al. 2015)
HMMER – used for searching sequence databases for homologs of protein sequences and for
making protein sequence alignments. It implements methods using probabilistic models called
profile hidden Markov models (profile HMMs) (Finn et al. 2015)
InterProScan (iprscan) – scans sequences for matches against the InterPro collection of protein
signature databases (Zdobnov and Apweiler 2001)
LINKS – long interval nucleotide k-mer scaffolder (Warren et al. 2015b)
MAKER2 – a portable and easily configurable genome annotation pipeline (Cantarel et al. 2008;
Holt and Yandell 2011)
MGEScan-LTR – identifies long terminal repeats (LTR) (Rho et al. 2007)
Minimap2 – a versatile pairwise aligner for genomic and spliced nucleotide sequences (Li 2016)
Miniasm – ultrafast de novo assembly for long noisy reads (though having no consensus step)
(Li 2016)
MCScanX – algorithm for detection of synteny and collinearity (Wang et al. 2012b)
MUMmer – a system for rapidly aligning entire genomes, whether in complete or draft form
(Marcais et al. 2018)
MUSCLE – multiple sequence alignment (Edgar 2004)
NextClip – a tool for comprehensive quality analysis and read preparation for Nextera Long Mate
Pair (LMP) libraries (Leggett et al. 2014)
Nseg – used to mask nucleic acid sequences, needed by RepeatScout
OMSSA – open mass spectrometry search algorithm (Geer et al. 2004)
OrthoCluster – synteny block identification (Ng et al. 2009)
Picard – a set of tools (in Java) for working with next-generation sequencing data in the BAM
(http://www.htslib.org/) format
Pilon – an automated genome assembly improvement and variant detection tool (Walker et al.
2014)
ProQuant – tool for label-free protein quantification of SEQUEST output for the Microsoft
Windows environment (Bridges et al. 2007)
a
Proteogenomic mapping tool – provides experimentally based structural annotations at a complete
genome level (Sanders et al. 2011)
a
ProtIDer – enhances proteomics based on EST or EST assemblies by generating databases of
matching highly homologous proteins (McCarthy et al. 2006)
a
(continued)
150
D. G. Peterson and M. Arick
