algorithm, GATK HaplotypeCaller [4], a number of steps are followed to determine
genotype likelihoods: First, regions of the genome where there is evidence of a variant
are defined (“SNP calling”), then the sequencing reads are used to identify the most likely
genotypes supported by these data, and each read is re-aligned to all these most likely
haplotypes in order to obtain the likelihoods per haplotype and per variant given the read
data. These are then input into Bayes’ formula to identify the most likely genotype for a
sample [4].
10.2.3 Heuristic Variant Calling
Other methods that do not rely on naive or Bayesian approaches have been developed;
these methods rely on heuristic quantities to call a variant site, such as a minimum coverage
and alignment quality thresholds, and stringent cut-offs like a minimum number of reads
supporting a variant allele. If performing somatic variant calling, a statistical test such as a
Fisher’s exact comparing the number of reference and alternate alleles in the tumor and
normal samples is then performed in order to determine the genotype at a site [7].
Parameters used for variant calling can be tuned, and generally this method will work
well with high-coverage sequencing data, but may not achieve an optimal equilibrium
between high specificity and high sensitivity at low to medium sequencing depths, or when
searching for low-frequency variants in a population [8].
The following GATK commands depict an example workflow for calling variants in
NGS data. The installation instruction is covered in Chap. 5.
First you can check all available GATK tools by typing
. If not already done,
you also have to install all required software packages (http://gatkforums.broadinstitute.
org/gatk/discussion/7098/howto-install-software-for-gatk-workshops) for GATK analyses
workflows. Moreover, be sure to set the PATH in your .bashrc to your GATK executable
PATH.
The GATK (v4) uses two files to access and safety check access to the reference files: a .
dict dictionary of the contig names and sizes and a .fai fasta index file to allow efficient
random access to the reference bases. You have to generate these files in order to be able to
use a Fasta file as reference.
128
P. Basurto-Lozada et al.
genotype likelihoods: First, regions of the genome where there is evidence of a variant
are defined (“SNP calling”), then the sequencing reads are used to identify the most likely
genotypes supported by these data, and each read is re-aligned to all these most likely
haplotypes in order to obtain the likelihoods per haplotype and per variant given the read
data. These are then input into Bayes’ formula to identify the most likely genotype for a
sample [4].
10.2.3 Heuristic Variant Calling
Other methods that do not rely on naive or Bayesian approaches have been developed;
these methods rely on heuristic quantities to call a variant site, such as a minimum coverage
and alignment quality thresholds, and stringent cut-offs like a minimum number of reads
supporting a variant allele. If performing somatic variant calling, a statistical test such as a
Fisher’s exact comparing the number of reference and alternate alleles in the tumor and
normal samples is then performed in order to determine the genotype at a site [7].
Parameters used for variant calling can be tuned, and generally this method will work
well with high-coverage sequencing data, but may not achieve an optimal equilibrium
between high specificity and high sensitivity at low to medium sequencing depths, or when
searching for low-frequency variants in a population [8].
The following GATK commands depict an example workflow for calling variants in
NGS data. The installation instruction is covered in Chap. 5.
First you can check all available GATK tools by typing
. If not already done,
you also have to install all required software packages (http://gatkforums.broadinstitute.
org/gatk/discussion/7098/howto-install-software-for-gatk-workshops) for GATK analyses
workflows. Moreover, be sure to set the PATH in your .bashrc to your GATK executable
PATH.
The GATK (v4) uses two files to access and safety check access to the reference files: a .
dict dictionary of the contig names and sizes and a .fai fasta index file to allow efficient
random access to the reference bases. You have to generate these files in order to be able to
use a Fasta file as reference.
128
P. Basurto-Lozada et al.
