querying in the context of their own biological questions [2]. Such
large data sets may not provide a complete understanding of a given
system, but they can be leveraged to help plan experiments or
generate hypotheses in silico. These hypotheses can then be rapidly
tested in the lab with the wide range of molecular techniques and
genetic resources currently available. This chapter aims to provide
an updated overview of web-based tools for querying data sets
generated by researchers, often funded by the National Science
Foundation Arabidopsis 2010 project in the U.S., whose stated
goal was to identify the functions of 25,000 genes in Arabidopsis
by 2010 [2], and by the AtGenExpress Consortium, an international effort to measure the Arabidopsis transcriptome under many
conditions and in different tissues.
Here, we will emphasize well-cited web-based tools that integrate data from several repositories—tools that draw from many
sources are often more useful to the typical Arabidopsis researcher
than single data source lab-based websites. Additionally, we would
like to refer you to a new unit on The Arabidopsis Information
Resource (TAIR at http://www.arabidopsis.org) published
recently by Eva Huala and colleagues that covers this excellent
sequence-centric Arabidopsis database [3]. The SIGnAL website
at http://signal.salk.edu/ [4] and https://www.araport.org [5] are
two further websites for exploring sequences and identifying insertions—we will touch on these briefly, along with websites for the
1001 Arabidopsis genomes project.
The main focus of the chapter will be on tools for exploring
transcriptome data sets, which are the most comprehensive of all
the large data types, and highlight the ones used for querying these
data sets both in targeted and correlative ways. Such tools can be
highly valuable for focusing the search for mutant phenotypes, or
for providing leads on novel genetic associations with a given
biological process, respectively. We will also examine several
resources for exploring protein–protein interactions in Arabidopsis,
or for performing promoter analyses. Integrating different data
types to improve function prediction is key to extracting even
more knowledge from these data sets.
As in the first edition of this chapter, we will use ABSCISIC
ACID INSENSITIVE 3, At3g24650 [6] as our “gene of interest.”
While this gene has long been known to be involved in seed
biology, we will hypothesize some additional functions using the
tools described here, some easily inferred at the cost of only a click
of the mouse. The programs and websites that will be discussed in
this chapter are listed in Table 1 in the Materials section. Two
further informative, if slightly older, review articles in the context
of bioinformatic tools for hypothesis generation are by Brady and
Provart [7] and by Usadel and colleagues [8]. We would also like to
point you to a recent article on the future of Arabidopsis informatics resources [9]. There, the International Arabidopsis Informatics
26
G. Alex Mason et al.
Précédent

- 36/947

Suivant