• Filename: The original filename of the file which was analyzed.
• File type: Says whether the file appeared to contain actual base calls or colorspace data
which had to be converted to base calls.
• Encoding: Says which ASCII encoding of quality values was found in this file.
• Total Sequences: A count of the total number of sequences processed. There are two
values reported, actual and estimated.
• Filtered Sequences: If running in Casava mode sequences flagged to be filtered will be
removed from all analyses. The number of such sequences removed will be reported
here. The total sequences count above will not include these filtered sequences and will
the number of sequences actually used for the rest of the analysis.
• Sequence Length: Provides the length of the shortest and longest sequence (long-read
sequencing) in the set. If all sequences are the same length, only one value is reported
(short-read sequencing).
• %GC: The overall %GC of all bases in all sequences.
7.3.1.2 Per Base Sequence Quality
This box-and-whisker plot shows the range of quality values (Phred-Scores; see Sect.
7.2.3) across all bases at each position in the FASTQ file.
The central red line is the median value, the yellow box represents the inter-quartile
range (25–75%), the upper and lower whiskers represent the 10% and 90% points, and the
blue line represents the mean quality.
On the x-axis the bases 1–10 are reported individually, then the bases are summarized in
bins. The number of base positions binned together depends on the length of the read, thus
shorter reads will have smaller windows and longer reads larger windows. The y-axis
depicts the Phred-Scores.
Often you will see a decreasing quality with increasing base position (Figs. 7.6 and 7.7).
This effect lies in the sequencing by synthesis technology of Illumina and is called phasing.
During each sequencing cycle chemicals that include variants for all four nucleotides are
washed over the flow cell. The nucleotides have a terminator cap so that only 1 base gets
incorporated. After the detection of the fluorescence signal the terminator cap is removed
and the next cycle can start. Accordingly, a synchronous sequencing of DNA fragments in
each cluster by expressing specific fluorescence signals is guaranteed (see Sect. 4.2). The
main reason for the decreasing sequence quality is that the blocker of a nucleotide is not
correctly removed after signal detection (phasing) and thus lead to light pollution during
signal detection. This error occurs more often over time and thus with an increasing read
length.
7.3.1.3 Per Tile Sequence Quality
The Per tile Sequence Quality Graph graph only appears in your FastQC report if you are
using an Illumina library. The original sequence identifiers are retained encoding the
flowcell tile from which each read came (Fig. 7.8). Reasons for seeing errors on this plot
could be transient problems such as bubbles going through the flow cell, or there could be
94
M. Kappelmann-Fenzl
• File type: Says whether the file appeared to contain actual base calls or colorspace data
which had to be converted to base calls.
• Encoding: Says which ASCII encoding of quality values was found in this file.
• Total Sequences: A count of the total number of sequences processed. There are two
values reported, actual and estimated.
• Filtered Sequences: If running in Casava mode sequences flagged to be filtered will be
removed from all analyses. The number of such sequences removed will be reported
here. The total sequences count above will not include these filtered sequences and will
the number of sequences actually used for the rest of the analysis.
• Sequence Length: Provides the length of the shortest and longest sequence (long-read
sequencing) in the set. If all sequences are the same length, only one value is reported
(short-read sequencing).
• %GC: The overall %GC of all bases in all sequences.
7.3.1.2 Per Base Sequence Quality
This box-and-whisker plot shows the range of quality values (Phred-Scores; see Sect.
7.2.3) across all bases at each position in the FASTQ file.
The central red line is the median value, the yellow box represents the inter-quartile
range (25–75%), the upper and lower whiskers represent the 10% and 90% points, and the
blue line represents the mean quality.
On the x-axis the bases 1–10 are reported individually, then the bases are summarized in
bins. The number of base positions binned together depends on the length of the read, thus
shorter reads will have smaller windows and longer reads larger windows. The y-axis
depicts the Phred-Scores.
Often you will see a decreasing quality with increasing base position (Figs. 7.6 and 7.7).
This effect lies in the sequencing by synthesis technology of Illumina and is called phasing.
During each sequencing cycle chemicals that include variants for all four nucleotides are
washed over the flow cell. The nucleotides have a terminator cap so that only 1 base gets
incorporated. After the detection of the fluorescence signal the terminator cap is removed
and the next cycle can start. Accordingly, a synchronous sequencing of DNA fragments in
each cluster by expressing specific fluorescence signals is guaranteed (see Sect. 4.2). The
main reason for the decreasing sequence quality is that the blocker of a nucleotide is not
correctly removed after signal detection (phasing) and thus lead to light pollution during
signal detection. This error occurs more often over time and thus with an increasing read
length.
7.3.1.3 Per Tile Sequence Quality
The Per tile Sequence Quality Graph graph only appears in your FastQC report if you are
using an Illumina library. The original sequence identifiers are retained encoding the
flowcell tile from which each read came (Fig. 7.8). Reasons for seeing errors on this plot
could be transient problems such as bubbles going through the flow cell, or there could be
94
M. Kappelmann-Fenzl
