gefunden werden.). For gene expression analysis, the sequencing reads are mapped on the
assembled transcriptome further leading to annotation and functional analysis [50]. The
biggest advantage of de novo assembly is its non-dependence on a reference genome. It is
advantageous in analyzing data for sequences originating from un-sequenced genomes,
partially annotated, unfinished genome drafts, and new/rare genomes including that of
epidemic/pandemic viruses such as Ebola, Nipah, and Covid-19. In most cases, the RNASeq analysis along with a de novo assembly provides first-hand information on transcripts
and phylogeny. On the contrary, performing a de novo assembly in conjugation with a
reference-based alignment could also help detecting novel transcripts and isoforms [51,
52]. Importantly, identification of novel splicing sites or transcripts via a de novo assembly
does not essentially require alignment information of pre-existing splice sites. Another
important advantage of a de novo assembly is its compatibility to both, short read and long
read sequencing platforms in contrast to a reference-based alignment that prefers the
former.
There are several popular de novo assembly tools available that include Rnnotator [53],
Trans-ABySS [54], and Trinity [55]. Besides, alternative and efficient approaches are
hybrid methods which use both, reference-based alignments and de novo assemblies [52]
((PMID: 29643938) Fehler! Verweisquelle konnte nicht gefunden werden., right).
11.6.2.1 Choice of de novo Assembly Tools
Currently, de Bruijn graphs are the most widely used algorithm for performing a de novo
assembly [56]. Hence, most of the popular de novo or reference-free assembly tools for
RNA-Seq data utilized de Bruijn graphs, including Velvet, Oases, and Trinity. Velvet and
Oases are used simultaneously [57, 58]. Velvet is a genome assembly tool that generates
assembly graphs that are further analyzed by Oases for finding paths in the graphs to
identify transcript isoforms and to generate a draft assembly. Trinity is based on three main
modules that perform an initial assembly and clustering, create individual de Bruijn graphs
for each cluster, and finally extract sequences representing transcript isoforms present at
enlisted gene locus [55]. Next, we describe how to use both Velvet/Oases and Trinity [49].
Velvet and Oases: Velvet is a genome assembler that calculates k-mers of data and
assigns the contigs into a de Bruijn graph. Similarly, Oases performs transcript assemblies
utilizing the output of Velvet. It segments the graphs generated by Velvet into transcript
isoforms linked to each locus. Both tools process single-end reads as default; however,
Velvet supports paired-end reads using a single file containing adjacently located read pair.
The main steps in a de novo assembly using Velvet and Oases are described below:
Creating de Bruijn graph The input data format is defined (FASTA/FASTQ; single end/
paired end) and the data is clustered based on a particular k-mer length (e.g., 20). A hash
table is created which is utilized to generate a de Bruijn graph for the defined k-mer size. In
this step, a hash table is created with defined k-mer size and then graph traversal is done to
create de Bruijn graphs by Velvet as follows:
11 Design and Analysis of RNA Sequencing Data
157
Précédent

- 165/225

Suivant