Contig Generation First Inchworm extracts all overlapping k-mers from the RNA-Seq
reads. Second, each unique k-mer is examined in decreasing order of abundance and
transcript contigs are generated using a greedy extension algorithm based on (k-1)-mer
overlaps. Finally, unique portions of alternatively spliced transcripts are saved for the next
step. Trinity is executed from the command line using a single Perl script “Trinity.pl”:
Description of parameters:
--seqType
the input sequence in FASTQ format (can be fa, or fq)
--max_memory
the suggested max memory to be use by Trinity in Gb of RAM
# if single reads:
--single
single reads, one or more file names, comma-delimited
# if paired reads:
--leftleft reads, one or more file names (separated by commas)
--right
right reads, one or more file names (separated by commas)
Creating de Bruijn Graph In the next step, the generated contigs are clustered using
Chrysalis based on regions originating from alternatively spliced transcripts or closely
related gene families. Following this, a de Bruijn graph for each cluster is created and the
reads are partitions among these contig clusters. These contig clusters are termed as
“components” and the partitioning of RNA-Seq reads into ‘components’ helps to process
large sets of reads in parallel.
Output In the final step, Butterfly processes individual de Bruijn graphs in parallel by
tracing RNA-Seq reads through each graph and determining connectivity based on the read
sequence. This results in reconstructed transcript sequences for alternatively spliced
isoforms along with transcripts that correspond to paralogous genes. The final output is a
single FASTA file containing reconstructed transcript sequences.
11 Design and Analysis of RNA Sequencing Data
159
Précédent

- 167/225

Suivant