book (Chaps. 4, 8 and 9). Base calling and read alignment results rely on the sequencing
instrument and algorithm used; therefore, it is important to state the confidence we have on
the assignment of each base. This is expressed by standard Phred quality scores [1].
QPhred ¼ À10 log 10 P error
ð
Þ
This measurement is an intuitive number that tells us the probability of the base call or
the alignment being wrong, and the higher Q is, the more confidence we have that there has
not been an error. For example, if Q ¼ 20, then that means there is a 1 in 100 chance of the
call or alignment being wrong, whereas if it is 30 then there is a 1 in 1000 chance of a
mistake. Subsequently, steps such as duplicate read marking and base call quality score
recalibration can be performed (See Chap. 7, Sect. 7.2.3).
Review Question 1
What would be the value of Q for a variant that has a 1 in 3000 chance of being wrong?
After the previous steps have taken place, and an alignment file has been produced
(usually in the BAM and CRAM file formats), the next step is to identify differences
between the reference genome and the genome that has been sequenced. To this effect,
there are different strategies that a researcher can use depending on their experiment, for
example, for germline analyses they might use algorithms that assume that the organism of
interest is diploid (or another, fixed ploidy) and for cancer genomes they may need to use
more flexible programs due to the presence of polyploidy and aneuploidy. In this Chapter,
we will focus on the former analyses, but the reader is referred to the publications on
somatic variant callers in the “Further Reading” section below if they want to learn more.
When identifying variants, and particularly if a researcher is performing whole exome or
genome sequencing, the main objective is to determine the genotype of the sample under
study at each position of the genome. For each variant position there will be a reference (R)
and an alternate (A) allele, the former refers to the sequence present in the reference
genome. Therefore, it follows that in the case of diploid organisms, there will be three
different possible genotypes: RR (homozygous reference), RA (heterozygous), and AA
(homozygous alternative).
10.2.1 Naive Variant Calling
A naive approach to determining these genotypes from a pile of sequencing reads mapped
to a site in the genome may be to count the number of reads with the reference and alternate
alleles and to establish hard genotype thresholds; for example, if more than 75% of reads
are R, then the genotype is called as RR; if these are less than 25%, then the genotype is
called as AA; and anything in between is deemed RA. However, even if careful steps are
taken to ensure that only high-quality bases and reads are counted in the calculation, this
method is still prone to under-calling variants in low-coverage data, as the counts from
10 Identification of Genetic Variants and de novo Mutations Based on NGS
125
Précédent

- 133/225

Suivant