>ps<-na.omit(read.table('Ath_glucose_germrate_emmax.ps',sep='\t'))
>colnames(ps)<-c('SNP','beta','pvalue')
>library(stringr) #no need to install stringr
>sep_ps<-str_split_fixed(ps$SNP,'_',2)
>chr<-as.numeric(sep_ps[,1])
>base<-as.numeric(sep_ps[,2])
>dat<-as.data.frame(cbind(chr,base,ps$pvalue))
>dat<-as.data.frame(cbind(chr,base,ps$pvalue))
>colnames(dat)<-c('CHR','BP','PVALUE')
>write.table(dat,file='Ath_glucost_germrate_aragwab_input.txt',sep='\t', row.names=F)
4. Upload genes that are known to be associated with the phenotype based on prior knowledge. These genes are used to find
the optimal parameters for araGWAB results. Input gene names
need to have TAIR10 locus ids (ATXGXXXXX, X is an integer),
and the file should be a simple text file with the extension “.txt”
to upload the file. We can also type user-input gene names in
TAIR10 locus ids separated by tabs, spaces, or commas.
5. There are several parameters for running araGWAB (Table 5).
We recommend using the default parameters since they were
optimized based on various GWAS datasets. If you cannot
retrieve genes of interest by boosting with default parameters,
try to increase the distance range.
Table 4
araGWAB input file format
chr
pos
p-value
1
1,051,029
0.0089792
5
3312
0.0001251
C
99,102
4.213000e–05
The input file format is a tab-separated file, with the first line being the name of three
columns: chromosome, base position, and p-value. araGWAB allows characters “C”
(chlorophyll) and “M” (mitochondria) in the first column
Genome-Wide Association Studies in Arabidopsis
205
Précédent

- 209/947

Suivant