You can have a look at a generated VCF and source file (i.e. *.bam) in the IGV Browser
[4–6]. You can find further information on VCF files in Chap. 10 and at https://samtools.
github.io/hts-specs/VCFv4.2.pdf.
7.2.10 SRA (Sequence Read Archive)
The Sequence Read Archive (SRA) is a bioinformatics database from NCBI (National
Center for Biotechnology Information) that provides a public repository for sequencing
data, generated by high-throughput sequencing. The SRA-toolkit (https://github.com/ncbi/
sra-tools) is needed to download the data [11].
The following command can be used to download an SRA file of a paired-end
sequencing experiment and to store mate I in *_I.fastq and mate II in *_II.fastq. The –
gzip option is used to minimize the size of the two fastq files.
In terms of a single-end sequencing experiment you would type:
Review Question 4
Find the sample with the accession SRR10257831 from the Sequence Read Archive and
find out the following information:
• Which species?
• Genome or transcriptome?
• What sequencing platform was used?
• What read length?
• Was it a single-end or paired-end sequencing approach?
7.3
Quality Check and Preprocessing of NGS Data
7.3.1 Quality Check via FastQC
FASTQ files (see Sect. 7.2.3) are the “raw data files” of any sequencing application, that
means they are “untouched.” Thus, this file format is used for Quality Check of sequencing
reads. The Quality Check procedure is commonly done with the FastQC tool written by
Simon Andrews of Babraham Bioinformatics (https://www.bioinformatics.babraham.ac.
uk/projects/fastqc/). FastQC and other similar tools are useful for assessing the overall
92
M. Kappelmann-Fenzl
[4–6]. You can find further information on VCF files in Chap. 10 and at https://samtools.
github.io/hts-specs/VCFv4.2.pdf.
7.2.10 SRA (Sequence Read Archive)
The Sequence Read Archive (SRA) is a bioinformatics database from NCBI (National
Center for Biotechnology Information) that provides a public repository for sequencing
data, generated by high-throughput sequencing. The SRA-toolkit (https://github.com/ncbi/
sra-tools) is needed to download the data [11].
The following command can be used to download an SRA file of a paired-end
sequencing experiment and to store mate I in *_I.fastq and mate II in *_II.fastq. The –
gzip option is used to minimize the size of the two fastq files.
In terms of a single-end sequencing experiment you would type:
Review Question 4
Find the sample with the accession SRR10257831 from the Sequence Read Archive and
find out the following information:
• Which species?
• Genome or transcriptome?
• What sequencing platform was used?
• What read length?
• Was it a single-end or paired-end sequencing approach?
7.3
Quality Check and Preprocessing of NGS Data
7.3.1 Quality Check via FastQC
FASTQ files (see Sect. 7.2.3) are the “raw data files” of any sequencing application, that
means they are “untouched.” Thus, this file format is used for Quality Check of sequencing
reads. The Quality Check procedure is commonly done with the FastQC tool written by
Simon Andrews of Babraham Bioinformatics (https://www.bioinformatics.babraham.ac.
uk/projects/fastqc/). FastQC and other similar tools are useful for assessing the overall
92
M. Kappelmann-Fenzl
