What You Will Learn in This Chapter
Next-generation sequencing experiments produce millions of short reads per sample
and the processing of those raw reads and their conversion into other file formats lead
to additional information on the obtained data. Various file formats are in use in order
to store and manipulate this information. This chapter presents an overview of the file
formats FASTQ, FASTA, SAM/BAM, GFF/GTF, BED, and VCF that are commonly used in analysis of next-generation sequencing data. Moreover, the structure
and function of the different file formats are reviewed. This chapter explains how
different file formats can be interpreted and what information can be gained from
their analysis.
7.1
Introduction
Analyzing NGS data means to handle really big data. Thus, one of the most important
things is to store all these “big data” in appropriate data formats to make them manageable.
The different data formats store in many cases the same type "of information, but not all
data formats are suitable for all bioinformatic tools. Always keep in mind that bioinformatic
software tools do not transform data into answers, they transform data into other data
formats. The answers result from investigating, understanding, and interpreting the various
data formats.
7.2
File Formats
7.2.1 Basic Notations
• Fragment: The molecule to be sequenced.
• Read: One sequenced part of a biological fragment (mate I or mate II).
• Mate I: The sequence of the 5’end of a paired-end sequencing approach.
• Mate II: The sequence of the 3’end of a paired-end sequencing approach.
• Sequencing depth: The number of all the sequences, reads, or bases represented in a
single sequencing experiment divided by the target region.
• Sequencing Coverage:
The theoretical redundancy of coverage (c) is described as LN/G, where L is the read
length, N is the number of reads, and G is the haploid genome length [1].
Sequencing coverage can be calculated in different ways depending on reference
points (whole genome, one locus, or one position in the genome):
1. One locus: # of bases mapping to the locus/size of locus.
80
M. Kappelmann-Fenzl
Précédent

- 89/225

Suivant