triggered by a long fixation time [9]. Although these events are detectable only at a small
fraction of the reads aligning to a particular site, they can become important when
analyzing a pool of genomes sequenced at low frequency or when studying tumor samples
that could have subclonal mutations—furthermore confounded by the tendency of some of
these tumor types toward having more real C>T mutations [10]. Another example comes
from the observation that DNA oxidation can happen during the shearing step most NGS
protocols have implemented, and that this results in artifactual C>A mutations [11].
Ancient DNA and ctDNA can also suffer from these problems [12, 13]. Therefore, a
researcher needs to consider their sample origin and preparation protocol and undertake
post-processing filtering steps accordingly.
10.2.5 How to Choose an Appropriate Algorithm for Variant Calling?
In addition to considering the variant calling method (e.g., naive, probabilistic, or heuristic)
that an algorithm implements, a researcher also needs to consider the types of genetic
variants that they are interested in analyzing, perhaps having to run several programs at the
same time to obtain a comprehensive picture of the genetic variation in their samples.
Genetic variants are usually classified into several groups according to their
characteristics (Fig. 10.2):
Fig. 10.2 Classes of genetic variants. Genetic variants ranging from a single base change, to the
insertion or deletion of several bases can occur in a genome. Structural variants are more complex and
encompass larger sections of a genome: At the top, a reference sequence, in the second row, a large
deletion (blue region), in the third row, a large insertion (red section), in the fourth row, an inversion,
and in the fifth row, a duplication. This figure is based on one drawn by Petr Danecek for a teaching
presentation
130
P. Basurto-Lozada et al.
Précédent

- 138/225

Suivant