sequencing by another orthogonal methodology such as capillary sequencing. If, on the
other hand, they are analyzing a large cohort of individuals in order to describe patterns of
variation, then they will need to be much stricter quality filters.
10.6.1 Variant Annotation
Variant annotation can help researchers filter and prioritize functionally important variants
for further study. Several tools for functional annotation have been developed; some of
them are based on public databases and are limited to known variants, while others have
been developed for the annotation of novel SNPs.
Functional prediction of variants can be done through different approaches, from
sequence-based analysis to structural impact on proteins. Predicted effects of identified
variants can be assessed through tools such as Ensembl-VEP [29] and SnpEff [30]. On top
of the predicted consequences on protein function (e.g., whether a variant is missense, stopgain, frameshift-inducing, etc.), these tools can also perform annotations at the level of
allele frequency against public databases such as 1000 Genomes and GnomAD, whether
the variant has been seen before either in populations or somatically in cancer (dbSNP and
COSMIC annotations), whether it falls in an evolutionarily conserved site (GERP and
PolyPhen-2 scores), and whether it has been found to have clinical relevance (ClinVar
annotations), among others. Genomic region-based annotations can also be performed,
referring to genomic elements other than genes, such as predicted transcription factor
binding sites, predicted microRNA target sites, and predicted stable RNA secondary
structures [31]. All these annotations can aid a researcher to focus on those variants
predicted to be associated to their phenotype of interest.
However, these steps to identify variant candidates are only part of the story. As we
mentioned above, even if the variants are real and seem to have an effect on gene function,
this alone is not enough evidence to link the variant causally to a phenotype [32].
Researchers should be wary of any potential positive associations and should consider
alternate hypotheses before reporting their identified variants as causal (or they may be
publicly challenged, see, for example, [33, 34]).
10.6.2 Evaluating the Evidence Linking Variants Causally to Phenotypes
After these essential filtering and annotation steps have been performed, a researcher then
needs to assess the amount of evidence supporting the potential causality of a genetic
variant. The first line of evidence needs to be statistical: Assuming a candidate variant
exists, the first question would be, how likely would it be to obtain an equivalent result by
chance if any other gene were to be considered? For example, a 2007 study by Chiu and
collaborators assumed that two novel missense genetic variants in the CARD3 gene were
causal of familial hypertrophic cardiomyopathy [35]. They assumed causality based on
10 Identification of Genetic Variants and de novo Mutations Based on NGS
135
other hand, they are analyzing a large cohort of individuals in order to describe patterns of
variation, then they will need to be much stricter quality filters.
10.6.1 Variant Annotation
Variant annotation can help researchers filter and prioritize functionally important variants
for further study. Several tools for functional annotation have been developed; some of
them are based on public databases and are limited to known variants, while others have
been developed for the annotation of novel SNPs.
Functional prediction of variants can be done through different approaches, from
sequence-based analysis to structural impact on proteins. Predicted effects of identified
variants can be assessed through tools such as Ensembl-VEP [29] and SnpEff [30]. On top
of the predicted consequences on protein function (e.g., whether a variant is missense, stopgain, frameshift-inducing, etc.), these tools can also perform annotations at the level of
allele frequency against public databases such as 1000 Genomes and GnomAD, whether
the variant has been seen before either in populations or somatically in cancer (dbSNP and
COSMIC annotations), whether it falls in an evolutionarily conserved site (GERP and
PolyPhen-2 scores), and whether it has been found to have clinical relevance (ClinVar
annotations), among others. Genomic region-based annotations can also be performed,
referring to genomic elements other than genes, such as predicted transcription factor
binding sites, predicted microRNA target sites, and predicted stable RNA secondary
structures [31]. All these annotations can aid a researcher to focus on those variants
predicted to be associated to their phenotype of interest.
However, these steps to identify variant candidates are only part of the story. As we
mentioned above, even if the variants are real and seem to have an effect on gene function,
this alone is not enough evidence to link the variant causally to a phenotype [32].
Researchers should be wary of any potential positive associations and should consider
alternate hypotheses before reporting their identified variants as causal (or they may be
publicly challenged, see, for example, [33, 34]).
10.6.2 Evaluating the Evidence Linking Variants Causally to Phenotypes
After these essential filtering and annotation steps have been performed, a researcher then
needs to assess the amount of evidence supporting the potential causality of a genetic
variant. The first line of evidence needs to be statistical: Assuming a candidate variant
exists, the first question would be, how likely would it be to obtain an equivalent result by
chance if any other gene were to be considered? For example, a 2007 study by Chiu and
collaborators assumed that two novel missense genetic variants in the CARD3 gene were
causal of familial hypertrophic cardiomyopathy [35]. They assumed causality based on
10 Identification of Genetic Variants and de novo Mutations Based on NGS
135
