10.6.3 Variant Filtering and Visualization Programs . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .136
10.7 A Practical Example Workflow . . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . . . .. . .139
References ...... ...... ...... ...... ...... ...... ...... ...... ...... ....... ...... ...... ...... ...... ...... ...141
What You Will Learn in This Chapter
In this chapter, we will discuss an overview of the bioinformatic process for the
identification of genetic variants and de novo mutations in data recovered from NGS
applications. We will pinpoint critical steps, describe the theoretical basis of different
variant calling algorithms, describe data formats, and review the different filtering
criteria that can be undertaken to obtain a set of high-confidence mutations. We will
also go over crucial issues to take into account when analyzing NGS data, such as
tissue source or the choice of sequencing machine. We also discuss different
methodologies for analyzing these variants depending on study context, considering
population-wide and family-focused analyses. Finally, we also do an overview of
available software for variant filtering and genetic data visualization.
10.1 Introduction: Quick Recap of a Sequencing Experiment Design
As we have seen throughout this book, NGS applications give researchers an all-access
pass to the building information of all biological organisms. After establishing the
biological question to be pursued and once the organism of interest has been sequenced,
the first step is to align this information against a reference genome (see Chap. 8). This
reference genome should be one that is as biologically close as it can be to the subject of
interest—if there is no reference genome or the one available is not reliable, then a possible
option is to attempt to build one (See Box: Genome assembly). After read mapping
and alignment, and quality control, one or several variant callers will need to be run to
identify the variants present in the query sequence. Finally, depending on the original aim,
different post-processing and filtering steps may also need to be deployed to extract
meaningful information out of the experiment.
10.2 How Are Novel Genetic Variants Identified?
The correct identification of variants depends on having accurately performed base calling,
and read mapping and alignment previously. “Base calling” refers to the determination of
the identity of a nucleotide from the fluorescence intensity information outputted by the
sequencing instrument. Read mapping is the process of determining where a read originates
from, using the reference genome, and read alignment is the process of finding the exact
differences between the sequences. These topics have already been reviewed throughout this
124
P. Basurto-Lozada et al.
Précédent

- 132/225

Suivant