PCR, read errors among other artifacts, QC is a critical step
before further data processing [24, 25]. We use fastQC
(Table 3) to check the sequence quality. Clean fastq file will
be generated. Reads generated through sRNA seq may cover
adapters partially. Untrimmed adapters will compromise the
mapping accuracy [23]. Several tools can be used for this step
including fastx [26], cutadapt [27], skewer [28], trimmomatic
[29], etc. Fastx is developed earlier and still widely used tool to
trim adapters, especially for 3
0 adapters. This method has a fair
balance between sensitivity and specificity with a medium speed
[28]. Herein we remove 3
0 adapter by using fastx clipper to get
“example-rmadapt.fastq” (see Note 10).
$ fastx_clipper -a NNNNTGGAATTCTCGGGTGCCAAGG -Q 33 -c -l 18 -o /example-rmadapt.fastq -i /example.fastq -v;
3. tRNA and rRNA removal.
In total RNA, sRNAs only account for about 1% or less,
while rRNA and tRNA consist 90%. Even if 28S and 18S rRNA
can be largely excluded by size selection, we may still have
contaminations of 5S rRNA (~120 nt) and tRNA (~70 nt).
Remove the potentially contaminated tRNA and rRNA is an
important step. Here we use bowtie, a fast and efficient program to align short reads to large genomes, to remove tRNA
and rRNA by using bowtie to get “example_rmstructure-rmadapt.fastq”.
$ bowtie-build structure_RNA.fa structure_RNA
$ bowtie -v 0 -a structure_RNA -p 10 -q /aim_directory/
example-rmadapt.fastq --un /file_localized_directory/ example_rmstructure-rmadapt.fastq;
4. Extract data of reads number between 18 to 30 nt.
Most sRNAs range from 18 to 30 nt in length, after
removal of contamination of structural RNAs, further cleaning
of sRNA reads is required to keep the 18–30 nt sRNAs. Extract
data of reads number between 18 and 30 nt by awk command
and obtain “example_18-30 nt.fastq”.
$ awk ’BEGIN {OFS ¼ "\n"} {header ¼ $0 ; getline seq ;
getline qheader ; getline qseq ; if (length(seq) >¼18 && length
(seq) <¼ 30) {print header, seq, qheader, qseq}}’ <
example_rmstructure-rmadapt.fastq > example_18-30nt.fastq ;
5. Data alignment.
Alignment is to map reads back to reference genomes.
More than 60 alignment tools have been developed up to
date [30]. Among them, bowtie/bowtie2 shows an ultrafast
speed and a high-quality score [30]. In our pipeline, we use
bowtie to align reads to TAIR10 to generate “example.sam”
[31]. Sequence alignment map (sam) file is the output of
aligners.
248
Di Sun et al.
Précédent

- 251/947

Suivant