7.3.1.11 Overrepresented Sequences
In a random library it is expected that most sequences occur only once in the final set. This
module lists all sequences which appear more often than expected. Finding that a certain
sequence is overrepresented either means that it is highly biologically relevant, or that the
library is contaminated. For each detected overrepresented sequence, the program looks for
matches in a database of common contaminants and will report the best hit it finds. Often
adapter sequences are detected as overrepresented reads. This can occur if you use a longread length and some of the library inserts are shorter than the read length resulting in readthrough to the adapter at the 3’ end of the read. How to remove adapter sequences is
described in Sect. 7.3.2.
7.3.2 Preprocessing of NGS Data-Adapter Clipping
If you detected an adapter contamination in your FastQC output file, it is highly
recommended to remove the adapter sequences from your sequences in the fastq file.
Therefore, the package cutadapt can be used, which we have already installed by
miniconda3. First you have to find the correct adapter sequence from the “Illumina
Costumer Sequence Letter” (https://support.illumina.com/content/dam/illumina-support/
documents/documentation/chemistry_documentation/experiment-design/illumina-adaptersequences-1000000002694-11.pdf).
The basic usage of cutadapt is:
For paired-end reads:
Moreover, you can also perform quality trimming with sequences depicting a bad phred
score. This can be done with the additional option
to trim low-quality bases
from 5’ and/or 3’ ends of each read before adapter removal. Applied to both reads if data is
paired. If one value is given, only the 3’ end is trimmed. If two comma-separated cutoffs
are given, the 5’ end is trimmed with the first cutoff, the 3’ end with the second.
Find some more options in terms of the command cutadapt by
.
Another tool for adapter trimming or removal of low-quality bases is
.
102
M. Kappelmann-Fenzl
In a random library it is expected that most sequences occur only once in the final set. This
module lists all sequences which appear more often than expected. Finding that a certain
sequence is overrepresented either means that it is highly biologically relevant, or that the
library is contaminated. For each detected overrepresented sequence, the program looks for
matches in a database of common contaminants and will report the best hit it finds. Often
adapter sequences are detected as overrepresented reads. This can occur if you use a longread length and some of the library inserts are shorter than the read length resulting in readthrough to the adapter at the 3’ end of the read. How to remove adapter sequences is
described in Sect. 7.3.2.
7.3.2 Preprocessing of NGS Data-Adapter Clipping
If you detected an adapter contamination in your FastQC output file, it is highly
recommended to remove the adapter sequences from your sequences in the fastq file.
Therefore, the package cutadapt can be used, which we have already installed by
miniconda3. First you have to find the correct adapter sequence from the “Illumina
Costumer Sequence Letter” (https://support.illumina.com/content/dam/illumina-support/
documents/documentation/chemistry_documentation/experiment-design/illumina-adaptersequences-1000000002694-11.pdf).
The basic usage of cutadapt is:
For paired-end reads:
Moreover, you can also perform quality trimming with sequences depicting a bad phred
score. This can be done with the additional option
to trim low-quality bases
from 5’ and/or 3’ ends of each read before adapter removal. Applied to both reads if data is
paired. If one value is given, only the 3’ end is trimmed. If two comma-separated cutoffs
are given, the 5’ end is trimmed with the first cutoff, the 3’ end with the second.
Find some more options in terms of the command cutadapt by
.
Another tool for adapter trimming or removal of low-quality bases is
.
102
M. Kappelmann-Fenzl
