Table 2 A list of publicly available sequence processing, assembly, alignment, comparison, and
annotation tools used by the authors at the Institute for Genomics, Biocomputing & Biotechnology
(IGBB) at Mississippi State University
Program name hyperlinked to program download site – description (reference)
3D-DNA – 3D de novo assembly (3D DNA) pipeline, i.e., a Hi-C scaffolder (Dudchenko et al.
2017)
ABySS – de novo, parallel, paired-end sequence assembler that is designed for short reads. The
single-processor version is useful for assembling genomes up to 100 Mb in size. The parallel
version is implemented using MPI and is capable of assembling larger genomes (Simpson et al.
2009; Jackman et al. 2017)
ALLPATHS-LG – short-read genome assembler from the Computational Research and Development group at the Broad Institute (Gnerre et al. 2011)
ArrayIDer – maps EST and probe IDs to their corresponding public database accessions (van den
Berg et al. 2009)
a
Augustus – predicts genes in eukaryotic genomic sequences (Stanke et al. 2004)
BioConductor – tools for the analysis and comprehension of high-throughput genomic data
(Zhang et al. 2003)
BioPerl – Perl libraries for biology (Stajich et al. 2002)
Biopython – a set of freely available tools for biological computation (Cock et al. 2009)
BLAST+ À Basic Local Alignment Search Tool (Camacho et al. 2009)
BLAT– sequence alignment tool (Kent 2002)
Bowtie/Bowtie2 – an ultrafast and memory-efficient tool for aligning sequencing reads to long
reference sequences (Langmead 2010; Langmead and Salzberg 2012)
BUSCO – assessing genome assembly and annotation completeness with Benchmarking Universal Single-Copy Orthologs (Simao et al. 2015)
BWA – a software package for mapping low-divergent sequences against a large reference
genome (Li and Durbin 2009)
Canu – adaptation of the Celera assembler designed for high-noise single-molecule sequence
assembly (PacBio & Nanopore) (Koren et al. 2017)
CAP3 – DNA sequence assembly (Huang and Madan 1999)
Circos – a software package for visualizing data and information in a circular format (Krzywinski
et al. 2009)
Cufflinks – assembles transcripts, estimates their abundances, and tests for differential expression
and regulation in RNA-Seq data (Trapnell et al. 2012)
DISCOVAR de novo – new genome assembly tool made by makers of ALLPATHS-LG (Love
et al. 2016)
EA-utils – command-line tools for processing biological sequencing data, barcode
demultiplexing, adapter trimming, etc. (Aronesty 2013)
EMBOSS – suite of bioinformatics software (Rice et al. 2000)
Exonerate – a generic tool for pairwise sequence comparison (Slater and Birney 2005)
FastQC – a quality control tool for high-throughput sequence data (Babraham Bioinformatics
2016)
GATK – a wide variety of tools with a primary focus on variant discovery, genotyping, and data
quality assurance (McKenna et al. 2010)
GeneMark – a family of gene prediction programs (Isono et al. 1994)
GenomeScope – script that estimates genome heterozygosity, repeat content, and size from
sequencing reads using a k-mer-based statistical approach (Vurture et al. 2017)
(continued)
Sequencing Plant Genomes
149
Précédent

- 158/342

Suivant