whereas in the opposite case, there will be a number of groups of
accessions that lie very close to each other. If subpopulations are
detected in the PCA plot, resampling the accessions for GWAS
analysis should be considered.
3.4 Obtaining
Phenotype Data
Phenotypes can be either qualitative or quantitative. For GWAS, a
quantitative phenotype is recommended over a qualitative phenotype, because it generally provides stronger statistical power in
detecting SNPs associated with the phenotype [22]. In addition,
if possible, obtaining phenotype data as images offers huge advantages over other types of data, because image data on phenotypes
enable us to revisit and obtain other related phenotype data (see
Note 4). In this protocol, we used the germination rate of Arabidopsis seedlings in 5.5% glucose concentration as a phenotype, since
it is known to suppress germination in the Col-0 ecotype [23].
3.5 Running GWAS
3.5.1 EMMAX Installation
1. Download the latest version of EMMAX and unzip. It is precompiled, so no further installation is required.
2. Set the path on your system to use EMMAX in any directory by
adding the following line to your ~/.bashrc file with any text
editor. In this section, we will use the preprocessed data from
Subheading 3.1 as input.
export PATH="$PATH:/path/to/emmax_directory"
3. Restart your terminal to apply the new setting.
3.5.2 Create a Kinship
Matrix
The kinship matrix allows EMMAX to account for genetic relatedness among individuals. Make sure that you have .tfam and .tped
files in the same directory that you are working on.
$emmax-kin -v -h -s -d 10 1001genomes_snps_maf0.1
You will get a single output kinship matrix file: “1001genomes_snps_maf0.1.hIBS.kinf.”
3.5.3 Input Phenotype
Data
1. Download an input phenotype data file, glucose_germrate.txt,
for the test analysis from https://netbiolab.org/wiki/glucose_
germrate.txt.
2. The input phenotype data file has three columns separated by
tabs, as shown in Table 2.
3. Input phenotype data should have the same order of accessions
as in the .tfam file. The .tfam file is organized in incremental
order of the accessions.
Genome-Wide Association Studies in Arabidopsis
195
accessions that lie very close to each other. If subpopulations are
detected in the PCA plot, resampling the accessions for GWAS
analysis should be considered.
3.4 Obtaining
Phenotype Data
Phenotypes can be either qualitative or quantitative. For GWAS, a
quantitative phenotype is recommended over a qualitative phenotype, because it generally provides stronger statistical power in
detecting SNPs associated with the phenotype [22]. In addition,
if possible, obtaining phenotype data as images offers huge advantages over other types of data, because image data on phenotypes
enable us to revisit and obtain other related phenotype data (see
Note 4). In this protocol, we used the germination rate of Arabidopsis seedlings in 5.5% glucose concentration as a phenotype, since
it is known to suppress germination in the Col-0 ecotype [23].
3.5 Running GWAS
3.5.1 EMMAX Installation
1. Download the latest version of EMMAX and unzip. It is precompiled, so no further installation is required.
2. Set the path on your system to use EMMAX in any directory by
adding the following line to your ~/.bashrc file with any text
editor. In this section, we will use the preprocessed data from
Subheading 3.1 as input.
export PATH="$PATH:/path/to/emmax_directory"
3. Restart your terminal to apply the new setting.
3.5.2 Create a Kinship
Matrix
The kinship matrix allows EMMAX to account for genetic relatedness among individuals. Make sure that you have .tfam and .tped
files in the same directory that you are working on.
$emmax-kin -v -h -s -d 10 1001genomes_snps_maf0.1
You will get a single output kinship matrix file: “1001genomes_snps_maf0.1.hIBS.kinf.”
3.5.3 Input Phenotype
Data
1. Download an input phenotype data file, glucose_germrate.txt,
for the test analysis from https://netbiolab.org/wiki/glucose_
germrate.txt.
2. The input phenotype data file has three columns separated by
tabs, as shown in Table 2.
3. Input phenotype data should have the same order of accessions
as in the .tfam file. The .tfam file is organized in incremental
order of the accessions.
Genome-Wide Association Studies in Arabidopsis
195
