• Integrative Genomics Viewer (IGV) [36] A very popular, highly interactive tool that is
able to process large amounts of sequencing data in different formats and display read
alignments, read- and variant-level annotations, and information from other databases.
Website: http://software.broadinstitute.org/software/igv/
• Galaxy [37] Another highly popular, web-based platform that allows researchers to
perform reproducible analyses through a graphical interphase. Users can load files in the
FASTA, BAM, and VCF formats, among others, and perform data analysis and variant
filtering in an intuitive way. Website: https://usegalaxy.org/
• VCF/Plotein [38] This web-based, interactive tool allows researchers to load files in the
VCF format and interactively visualize and filter variants in protein-coding genes. It
incorporates annotations from other external databases.
Website: https://vcfplotein.liigh.unam.mx/
Take Home Message
• There are a number of different methods for performing variant calling, these can
be naive, probabilistic, and heuristic. Probabilistic methods are the most widely
used and implement a form of Bayes’ Theorem. However, algorithm choice will
depend on the researcher’s study design.
• Sample storage and preparation methods may introduce errors that increase false
positive calls and therefore should be considered when designing an analysis
pipeline.
• Post-variant calling filters that analyze the distribution of variants across all
sequencing reads will usually need to be applied to data in order to reduce false
positive calls.
• True de novo genetic variants can be identified by analyzing trios with an affected
child, in other scenarios a number of annotations and filtering steps need to be
applied to identify candidate variants.
• For a researcher to ascribe phenotype causality to a genetic variant, the result of
gene- and variant-level annotations are not enough; a number of further statistical,
bioinformatic, and functional considerations need to be taken into account.
• Variant filtering and visualization tools can aid a researcher to perform the above
mentioned steps in an easy and intuitive way.
Answers to Review Questions
Answer to Question 1: Q ¼ 34.77.
Answer to Question 2: The logical option is for the researcher to go back and analyze,
through tools such as Samtools depth, whether indeed there is enough coverage at every
assessed site, and to mark it as “no call” otherwise. A novel VCF format, called gVCF
and outputted by GATK, can now give reference call confidence scores.
Answer to Review Question 3: There are a number of filters already implemented in
variant filtering tools, some of these are a threshold for Phred-scaled variant quality,
10 Identification of Genetic Variants and de novo Mutations Based on NGS
137
able to process large amounts of sequencing data in different formats and display read
alignments, read- and variant-level annotations, and information from other databases.
Website: http://software.broadinstitute.org/software/igv/
• Galaxy [37] Another highly popular, web-based platform that allows researchers to
perform reproducible analyses through a graphical interphase. Users can load files in the
FASTA, BAM, and VCF formats, among others, and perform data analysis and variant
filtering in an intuitive way. Website: https://usegalaxy.org/
• VCF/Plotein [38] This web-based, interactive tool allows researchers to load files in the
VCF format and interactively visualize and filter variants in protein-coding genes. It
incorporates annotations from other external databases.
Website: https://vcfplotein.liigh.unam.mx/
Take Home Message
• There are a number of different methods for performing variant calling, these can
be naive, probabilistic, and heuristic. Probabilistic methods are the most widely
used and implement a form of Bayes’ Theorem. However, algorithm choice will
depend on the researcher’s study design.
• Sample storage and preparation methods may introduce errors that increase false
positive calls and therefore should be considered when designing an analysis
pipeline.
• Post-variant calling filters that analyze the distribution of variants across all
sequencing reads will usually need to be applied to data in order to reduce false
positive calls.
• True de novo genetic variants can be identified by analyzing trios with an affected
child, in other scenarios a number of annotations and filtering steps need to be
applied to identify candidate variants.
• For a researcher to ascribe phenotype causality to a genetic variant, the result of
gene- and variant-level annotations are not enough; a number of further statistical,
bioinformatic, and functional considerations need to be taken into account.
• Variant filtering and visualization tools can aid a researcher to perform the above
mentioned steps in an easy and intuitive way.
Answers to Review Questions
Answer to Question 1: Q ¼ 34.77.
Answer to Question 2: The logical option is for the researcher to go back and analyze,
through tools such as Samtools depth, whether indeed there is enough coverage at every
assessed site, and to mark it as “no call” otherwise. A novel VCF format, called gVCF
and outputted by GATK, can now give reference call confidence scores.
Answer to Review Question 3: There are a number of filters already implemented in
variant filtering tools, some of these are a threshold for Phred-scaled variant quality,
10 Identification of Genetic Variants and de novo Mutations Based on NGS
137
