associations [1]. Genome-wide association studies (GWAS) provide
unbiased and comprehensive information on candidate loci that
should be investigated more deeply [2]. The reference plant Arabidopsis thaliana is well suited for GWAS because it is possible to
maintain inbred lines for association mapping via continued selffertilization, enabling repeated phenotyping of mapping accessions
with identical genotypes [3]. The pilot study on Arabidopsis GWAS
was performed with 95 accessions, which was sufficient to identify
genes already known for complex traits [4]. A follow-up study used
up to 200 accessions to find the single nucleotide polymorphism
(SNP)-phenotype associations in 107 phenotypes [5]. This study
showed that even with the number of accessions as low as 96, significant SNP-phenotype associations can be identified. Subsequently, exploiting the 250K SNP array [6], 1307 natural
accessions were genotyped, providing a huge resource of genotype
data that further facilitated Arabidopsis GWAS [7]. Now, the 1001
Genomes Project [8] has been completed and provides a surplus of
genetic variation data from 1135 sequenced natural accessions.
With this rich resource in hand, it is possible to simply obtain
seeds for mapping accessions from the stock center and perform
GWAS to identify genomic loci associated with a phenotype of
interest.
To identify the associations between single nucleotide polymorphism (SNP) sites and complex phenotypes, exploiting mixed
linear models is a good practice, because they effectively control for
population structure [9]. However, since there are more than a
million SNPs for statistical tests, solving the mixed linear models in
a traditional maximum likelihood estimation would be computationally expensive. To overcome this limitation, Efficient MixedModel Association (EMMA) [10] was developed and further
enhanced to EMMA eXpedited (EMMAX) [11] and Population
Parameters Previously Determined (P3D) [12], which substantially
reduce the computing time from days to minutes.
Although GWAS is an effective method, the resultant candidate
loci may not be able to explain all the phenotypic outcomes, a
situation that is referred to as the “missing heritability” problem
[2]. The unexplainable genetic inheritance can be partially
addressed by identifying rare variants, allelic heterogeneity, genetic
heterogeneity, epistatic interactions, and epigenetic variation,
which requires sophisticated genotyping [2]. Some of the
phenotype-associated SNPs that are missed owing to insufficient
statistical power may be rescued by integrating GWAS signals with
other types of functional genomics data. Augmentation of GWAS
signals via incorporation of pathway or network information has
been found to be effective in A. thaliana [13–16] as well as in a
crop species [17].
188
Tak Lee and Insuk Lee
unbiased and comprehensive information on candidate loci that
should be investigated more deeply [2]. The reference plant Arabidopsis thaliana is well suited for GWAS because it is possible to
maintain inbred lines for association mapping via continued selffertilization, enabling repeated phenotyping of mapping accessions
with identical genotypes [3]. The pilot study on Arabidopsis GWAS
was performed with 95 accessions, which was sufficient to identify
genes already known for complex traits [4]. A follow-up study used
up to 200 accessions to find the single nucleotide polymorphism
(SNP)-phenotype associations in 107 phenotypes [5]. This study
showed that even with the number of accessions as low as 96, significant SNP-phenotype associations can be identified. Subsequently, exploiting the 250K SNP array [6], 1307 natural
accessions were genotyped, providing a huge resource of genotype
data that further facilitated Arabidopsis GWAS [7]. Now, the 1001
Genomes Project [8] has been completed and provides a surplus of
genetic variation data from 1135 sequenced natural accessions.
With this rich resource in hand, it is possible to simply obtain
seeds for mapping accessions from the stock center and perform
GWAS to identify genomic loci associated with a phenotype of
interest.
To identify the associations between single nucleotide polymorphism (SNP) sites and complex phenotypes, exploiting mixed
linear models is a good practice, because they effectively control for
population structure [9]. However, since there are more than a
million SNPs for statistical tests, solving the mixed linear models in
a traditional maximum likelihood estimation would be computationally expensive. To overcome this limitation, Efficient MixedModel Association (EMMA) [10] was developed and further
enhanced to EMMA eXpedited (EMMAX) [11] and Population
Parameters Previously Determined (P3D) [12], which substantially
reduce the computing time from days to minutes.
Although GWAS is an effective method, the resultant candidate
loci may not be able to explain all the phenotypic outcomes, a
situation that is referred to as the “missing heritability” problem
[2]. The unexplainable genetic inheritance can be partially
addressed by identifying rare variants, allelic heterogeneity, genetic
heterogeneity, epistatic interactions, and epigenetic variation,
which requires sophisticated genotyping [2]. Some of the
phenotype-associated SNPs that are missed owing to insufficient
statistical power may be rescued by integrating GWAS signals with
other types of functional genomics data. Augmentation of GWAS
signals via incorporation of pathway or network information has
been found to be effective in A. thaliana [13–16] as well as in a
crop species [17].
188
Tak Lee and Insuk Lee
