3.2 Analyses of
Nucleotide Diversity
and Environmental
Associations for Your
Favorite Genes
Laboratories interested in the genetic dissection of specific
biological processes, genes, or gene families, but not specialized
in the analysis of natural genetic variation, can now exploit current
genomic resources to explore the amount, pattern, and geographic
or environmental distribution, of Arabidopsis nucleotide diversity
in specific genes. These analyses might allow the identification of
natural mutations that can be of interest for functional analyses, and
to define hypotheses about the adaptive potential of particular
genes. To address this:
1. Generate a variant call format (vcf) or a fasta format file for the
aligned sequence of your gene using the 1001 genomes database tools (sections tools/vcf or tools/pseudogenomes; see
Table 1), and selecting the accessions and locus of interest (see
Notes 30–33).
2. Estimate the amount and pattern of nucleotide diversity using
specific software packages for nucleotide diversity analyses
(DnaSP, MEGA or TASSEL; see Table 1). This analysis will
inform if your gene shows more or less diversity than the
average of Arabidopsis genome [60]. Carry out sliding window
analyses to find if the distribution of nucleotide diversity across
the gene is homogeneous or there are high and/or low diversity regions, which might reflect differential evolutionary histories. Evaluate if the nucleotide diversity in your gene might
be selectively neutral using existing population genetic tests,
such as Tajima, or Fu and Li tests (see Notes 34 and 35).
3. Analyze the predicted functional effect of the single nucleotide
polymorphisms (SNPs) and insertion/deletion (indels) segregating in your gene to detect missense mutations in conserved
domains or potentially high impact polymorphisms (see Notes
36 and 37).
4. Construct a neighbor-joining (NJ) tree and carry out a principal component analysis (PCA) to determine the genetic relationships among accessions for your gene. These analyses will
identify the number of different genotypes (haplotypes), and if
there are multiple groups of weakly differentiated haplotypes
(haplogroups) that are highly differentiated for your gene.
Estimate the relative genetic differentiation between haplogroups and identify the polymorphisms contributing to it
(see Notes 30, 38, and 39).
5. Generate a haplotype file including all SNPs and haplotype
frequencies. Use this data file to reconstruct a maximum parsimony phylogenetic network, which will reflect the simplest
evolutionary path for your gene. When including A. lyrata
sequence, this analysis might also infer the potential
A. thaliana ancestral haplotype (see Notes 40–42).
98
Carlos Alonso-Blanco et al.
Précédent

- 106/947

Suivant