10.3 Working with Variants
Variant calling usually outputs VCF files (See Sect. 7.2 File Formats) [18]. To recap, VCF
files are plain text files that contain genotype information about all samples in a sequencing
project. A VCF file is arranged like a matrix, with chromosome positions in rows and
variant and sample information in columns. The sample information contains the genotype
called by the algorithm along with a wealth of information such as (depending on the
algorithm) genotype likelihoods and sequencing depth supporting each possible allele,
among others. For each variant position, the file also contains information outputted by the
variant caller such as the reference and alternate alleles, the Phred-based variant quality
score, whether overlapping variants in other sequencing or genotype projects have been
found, etc. Crucially, this file also contains a column called “FILTER,” where information
about whether further quality filters have been applied to the calls and which ones. We will
review here some of the most common filters that researchers should consider applying to
their data once it has already been called.
Review Question 2
How do you think a researcher can deal with the uncertainty about false negatives, i.e.
sites where a variant has not been called? How can they be sure there is no variation
there and it is not, let us say, a lack of sequence coverage?
10.4 Applying Post-variant Calling Filters
So far, we have seen a number of steps where a researcher must be careful to increase both
the sensitivity and specificity of their set of calls in order to have an accurate view of the
amount and types of sequencing variation present in their samples. However, there are also
a number of post-calling filtering steps that should be applied in the majority of cases in
order to further minimize the amount of false positive calls. Here we will review the main
such filters, but the reader is referred to [18] if they wish to delve deeper into this topic.
Strand Bias Filter . Sometimes a phenomenon, referred to as strand bias, can be observed
where one base is called only in reads in one direction, whereas it is absent in the reads in
the other direction. This is evidently an error introduced during the preceding steps, and can
be detected through a strand bias filter. This filter applies a Fisher’s exact test comparing
the number of forward and reverse reads with reference and alternate alleles, and if the Pvalue is sufficiently small as determined through a pre-chosen threshold, then the variant is
deemed an artifact.
Variant Distance and End-Distance Bias Filters These filters were primarily developed
to deal with RNA sequencing data when aligned to a reference genome [19]. Therefore, if a
variant is mostly or only supported by differences in the last bases of each read, the call may
132
P. Basurto-Lozada et al.
Variant calling usually outputs VCF files (See Sect. 7.2 File Formats) [18]. To recap, VCF
files are plain text files that contain genotype information about all samples in a sequencing
project. A VCF file is arranged like a matrix, with chromosome positions in rows and
variant and sample information in columns. The sample information contains the genotype
called by the algorithm along with a wealth of information such as (depending on the
algorithm) genotype likelihoods and sequencing depth supporting each possible allele,
among others. For each variant position, the file also contains information outputted by the
variant caller such as the reference and alternate alleles, the Phred-based variant quality
score, whether overlapping variants in other sequencing or genotype projects have been
found, etc. Crucially, this file also contains a column called “FILTER,” where information
about whether further quality filters have been applied to the calls and which ones. We will
review here some of the most common filters that researchers should consider applying to
their data once it has already been called.
Review Question 2
How do you think a researcher can deal with the uncertainty about false negatives, i.e.
sites where a variant has not been called? How can they be sure there is no variation
there and it is not, let us say, a lack of sequence coverage?
10.4 Applying Post-variant Calling Filters
So far, we have seen a number of steps where a researcher must be careful to increase both
the sensitivity and specificity of their set of calls in order to have an accurate view of the
amount and types of sequencing variation present in their samples. However, there are also
a number of post-calling filtering steps that should be applied in the majority of cases in
order to further minimize the amount of false positive calls. Here we will review the main
such filters, but the reader is referred to [18] if they wish to delve deeper into this topic.
Strand Bias Filter . Sometimes a phenomenon, referred to as strand bias, can be observed
where one base is called only in reads in one direction, whereas it is absent in the reads in
the other direction. This is evidently an error introduced during the preceding steps, and can
be detected through a strand bias filter. This filter applies a Fisher’s exact test comparing
the number of forward and reverse reads with reference and alternate alleles, and if the Pvalue is sufficiently small as determined through a pre-chosen threshold, then the variant is
deemed an artifact.
Variant Distance and End-Distance Bias Filters These filters were primarily developed
to deal with RNA sequencing data when aligned to a reference genome [19]. Therefore, if a
variant is mostly or only supported by differences in the last bases of each read, the call may
132
P. Basurto-Lozada et al.
