7.2.3 FASTQ
Once you have sequenced your samples of interest you get back your sequencing data in a
certain format storing the reads information. Each sequencing read (i.e., paired-end) is
structured as depicted in Fig. 7.1. All reads and their information are stored in a file format
called FASTQ.
Sequencing facilities often store the read information in *.fastq or unaligned *.bam files.
The latter can be transformed in a *.fastq file via BEDTools:
The FASTQ format is a text-based standard format for storing both, a DNA sequence
and its corresponding quality scores from NGS. There are four lines per sequencing read.
FASTQ format example:
 The first line starts with '@', followed by the label.
 The second line represents the sequence of the read.
 The third line starts with '+‘, serving as a separator.
 The fourth line contains the Q scores (quality values for sequence in line 2)
represented as ASCII characters
Fig. 7.1 Sequencing read (ends of DNA fragment for mate pairs)
82
M. Kappelmann-Fenzl
Précédent

- 91/225

Suivant