suitable to align millions of NGS reads. This led to the development of advanced
algorithms that can meet this task, allow distinguishing polymorphisms from
mutations and sequencing errors from true sequence deviations. For a basic understanding, the differences between global and local alignment and the underlying
algorithms are described in a simplified way in this chapter, as well as the main
difference between BLAST and NGS alignment is described in a simplified way in
this chapter. Moreover, different alignment tools and their basic usage are presented,
which enables the reader to perform and understand alignment processes of sequencing reads to any genome using the respective commands.
9.1
Introduction
Sequence Alignment is a crucial step of the downstream analysis of NGS data, where
millions of sequenced DNA fragments (reads) have to be aligned with a selected reference
sequence within a reasonable time. However, the problem here is to find the correct
position in the reference genome from where the read originates. Due to the repetitive
regions of the genome and the limited length of the reads ranging from 50 to 150 bp, it often
happens that shorter reads can map at several locations in the genome. On the other hand, a
certain degree of flexibility for differences to the reference genome must be allowed during
alignment in order to identify point mutations and other genetic changes.
Due to the massive amount of data generated during NGS analyses, all alignment
algorithms use additional data structures (indices) that allow fast access and matching of
sequence data. These indices are generated either over all generated reads or over the entire
reference genome, depending on the used algorithm. Algorithms from computer science like
hash tables or methods from data compression like suffix arrays are popularly implemented in
the alignment tools. With the help of these algorithms, it is possible, for example, to compare
over 100 GB of sequence data from NGS analyses with the human reference genome in just a
few hours. With the help of high parallelization of computing capacity (CPUs), it is possible
to reduce this time significantly. Thus, even large amounts of sequencing data from wholeexome or whole-genome sequencing can be efficiently processed.
9.2
Alignment Definition
The next step is the alignment of your sequencing reads to a reference genome or
transcriptome. The main problem you have to keep in mind is that the human genome is
really big and it is complex too. Sequencers are able to produce billions of reads per run and
are prone to errors. Thus, an accurate alignment is a time-consuming process.
112
M. Kappelmann-Fenzl
algorithms that can meet this task, allow distinguishing polymorphisms from
mutations and sequencing errors from true sequence deviations. For a basic understanding, the differences between global and local alignment and the underlying
algorithms are described in a simplified way in this chapter, as well as the main
difference between BLAST and NGS alignment is described in a simplified way in
this chapter. Moreover, different alignment tools and their basic usage are presented,
which enables the reader to perform and understand alignment processes of sequencing reads to any genome using the respective commands.
9.1
Introduction
Sequence Alignment is a crucial step of the downstream analysis of NGS data, where
millions of sequenced DNA fragments (reads) have to be aligned with a selected reference
sequence within a reasonable time. However, the problem here is to find the correct
position in the reference genome from where the read originates. Due to the repetitive
regions of the genome and the limited length of the reads ranging from 50 to 150 bp, it often
happens that shorter reads can map at several locations in the genome. On the other hand, a
certain degree of flexibility for differences to the reference genome must be allowed during
alignment in order to identify point mutations and other genetic changes.
Due to the massive amount of data generated during NGS analyses, all alignment
algorithms use additional data structures (indices) that allow fast access and matching of
sequence data. These indices are generated either over all generated reads or over the entire
reference genome, depending on the used algorithm. Algorithms from computer science like
hash tables or methods from data compression like suffix arrays are popularly implemented in
the alignment tools. With the help of these algorithms, it is possible, for example, to compare
over 100 GB of sequence data from NGS analyses with the human reference genome in just a
few hours. With the help of high parallelization of computing capacity (CPUs), it is possible
to reduce this time significantly. Thus, even large amounts of sequencing data from wholeexome or whole-genome sequencing can be efficiently processed.
9.2
Alignment Definition
The next step is the alignment of your sequencing reads to a reference genome or
transcriptome. The main problem you have to keep in mind is that the human genome is
really big and it is complex too. Sequencers are able to produce billions of reads per run and
are prone to errors. Thus, an accurate alignment is a time-consuming process.
112
M. Kappelmann-Fenzl
