length can of course change after trimming sequences from the end due to poor quality base
calls or adapter contamination (see Sect. 7.3.2). If you performed long-read sequencing,
you will obtain a distribution of various read lengths, which means some reads are shorter,
most reads have an enriched size distribution, and some are longer.
7.3.1.10 Sequence Duplication Levels
There are two sources of duplicate reads: PCR duplication due to biased PCR enrichment
or really overrepresented sequences such as very abundant transcripts in an RNA-Seq
library. PCR duplicates misrepresent the true proportion of sequences in your starting
material, whereas really overrepresented sequences do faithfully represent your input.
Thus, in DNA-Seq nearly 100% of your reads should be unique (Fig. 7.15), in RNA-Seq
duplicate reads of highly abundant transcripts will be observed, however the duplication is
normal in this case (Fig. 7.16).
100
90
80
70
60
50
40
30
20
10
0
1 2 3 4 5 6 7 8 9 11 13 15 17 19 21 23
Position in read (bp)
N content across all bases
%N
25 27 29 31 33 35 37 39 41 43 45 47 49
Fig. 7.14 Per base N content. (source: https://rtsf.natsci.msu.edu)
100
M. Kappelmann-Fenzl
Précédent

- 109/225

Suivant