9 Genomic Techniques and How to Apply Them to Marine Questions
367
9.4.2.8 Pure Analysis Systems
Pure analysis systems provide variable collections of analysis algorithms, but they
normally do not take care of data-management. Data is normally stored in files and
user interaction is possible via a graphical user interface (GUI). One of the first
examples was the Cluster and TreeView software, which was developed by Eisen
et al. (1998), and focused on clustering algorithms. Cluster and TreView work only
on Windows. The Genesis software is another application that provides various clustering methods (Sturn et al. 2002). It is implemented in Java and thus operating
system independent. ArrayNorm (Pieler et al. 2004) and MIDAS (Saeed et al. 2003)
are programs that provide normalization for two-colour arrays and some statistical
tests.
Users of Affymetrix
R
microarrays can use the dChip software that provides normalization, clustering, and classification methods for oligonucleotide arrays (Li and
Wong 2001, Lin et al. 2004).
In summary, pure analysis tools are recommended for small to medium size laboratories. In general, they have minimal resource requirements, but data-management
and collaborative functions are not provided. For larger projects with many participants these systems can be useful to complement a database-based system.
9.4.2.9 General Purpose Database Systems
There is also a class of systems that combine general data analysis functions with
data-management capabilities. These systems use a database system to systematically store and retrieve data. Furthermore, all allow the annotation of experiments
by storing protocols. These systems include the open source systems BASE (Saal
et al. 2002), MARS (Maurer et al. 2005), MADAM (Saeed et al. 2003), and EMMA
(Dondrup et al. 2003) (Dondrup et al. 2009). All systems provide user and account
management and collaborative functions such as making experimental data publicly
available. Often, the software is accessed via a web-browser. Unfortunately, these
systems have relatively high requirements with respect to installation and administration of the server, which make them impractical for small laboratories. For
larger institutions it might be a good option to set up and maintain a central installation of such a system, in particular if researchers work as part of an international
collaboration.
9.4.2.10 R and BioConductor
The statistical environment R (Team 2008) has a special and prominent position
among all analysis systems. It is a general purpose statistical environment and also
a powerful programming language. The BioConductor project bundles and provides
add-on packages for many bioinformatics applications including microarrays and
sequence analysis (Gentleman et al. 2004). In comparison to other analysis systems
R is highly flexible and provides the largest amount of analysis algorithms. General
statistical functions also applicable to microarrays comprise various statistical test
367
9.4.2.8 Pure Analysis Systems
Pure analysis systems provide variable collections of analysis algorithms, but they
normally do not take care of data-management. Data is normally stored in files and
user interaction is possible via a graphical user interface (GUI). One of the first
examples was the Cluster and TreeView software, which was developed by Eisen
et al. (1998), and focused on clustering algorithms. Cluster and TreView work only
on Windows. The Genesis software is another application that provides various clustering methods (Sturn et al. 2002). It is implemented in Java and thus operating
system independent. ArrayNorm (Pieler et al. 2004) and MIDAS (Saeed et al. 2003)
are programs that provide normalization for two-colour arrays and some statistical
tests.
Users of Affymetrix
R
microarrays can use the dChip software that provides normalization, clustering, and classification methods for oligonucleotide arrays (Li and
Wong 2001, Lin et al. 2004).
In summary, pure analysis tools are recommended for small to medium size laboratories. In general, they have minimal resource requirements, but data-management
and collaborative functions are not provided. For larger projects with many participants these systems can be useful to complement a database-based system.
9.4.2.9 General Purpose Database Systems
There is also a class of systems that combine general data analysis functions with
data-management capabilities. These systems use a database system to systematically store and retrieve data. Furthermore, all allow the annotation of experiments
by storing protocols. These systems include the open source systems BASE (Saal
et al. 2002), MARS (Maurer et al. 2005), MADAM (Saeed et al. 2003), and EMMA
(Dondrup et al. 2003) (Dondrup et al. 2009). All systems provide user and account
management and collaborative functions such as making experimental data publicly
available. Often, the software is accessed via a web-browser. Unfortunately, these
systems have relatively high requirements with respect to installation and administration of the server, which make them impractical for small laboratories. For
larger institutions it might be a good option to set up and maintain a central installation of such a system, in particular if researchers work as part of an international
collaboration.
9.4.2.10 R and BioConductor
The statistical environment R (Team 2008) has a special and prominent position
among all analysis systems. It is a general purpose statistical environment and also
a powerful programming language. The BioConductor project bundles and provides
add-on packages for many bioinformatics applications including microarrays and
sequence analysis (Gentleman et al. 2004). In comparison to other analysis systems
R is highly flexible and provides the largest amount of analysis algorithms. General
statistical functions also applicable to microarrays comprise various statistical test
