these reads will likely not reach the set threshold, relevant information can be ignored due
to hard quality filters, and also it would not give any measure of confidence [2] (Fig. 10.1).
Even though this type of method was used in early algorithms, it has been dropped in favor
of other algorithms that are able to deal with errors and low-coverage data better.
10.2.2 Bayesian Variant Calling
The sequencing of two mixed molecules of DNA is a probabilistic event, as well as the
occurrence of errors in previous steps of this process, and therefore, a more informative
approach would include information about the prior probability of a variant occurring at
that site and the amount of information supporting each of the potential genotypes. To
address this need, most currently used variant callers implement a Bayesian probabilistic
approach at their core in order to assign genotypes [3]. Examples of the most commonly
used algorithms using this method are GATK HaplotypeCaller [4] and
bcftools mpileup (Formerly known as Samtools mpileup) [5].
This probabilistic approach uses the widely known Bayes’ Theorem, which, in this
context, is able to express the posterior probability of a genotype given the sequencing data
Fig. 10.1 Naive variant calling. In this method, reads are aligned to the reference sequence (green)
and a threshold of the proportion of reads supporting each allele for calling genotypes is established
(top). Then, at each position, the proportion of reads supporting the alternative allele is calculated and,
based on the dosage of the alternative allele, a genotype is established. Yellow: a position where a
variant is present but the proportion of alternative alleles does not reach the threshold (1/6 < 0.25). In
light blue, positions where a variant has been called. This figure is based on one drawn by Petr
Danecek for a teaching presentation
126
P. Basurto-Lozada et al.
Précédent

- 134/225

Suivant