9 Genomic Techniques and How to Apply Them to Marine Questions
365
not be the case, for example, the Wilcoxon’s Rank Sum Statistic, or the Significance
Analysis of Microarrays (SAM) (Tusher et al. 2001) method. Other methods, such as
the CyberT method (Baldi and Long 2001), attempt to address the problem of small
numbers of replicates. The principal of CyberT is to borrow variance information
from other genes on the microarray having a similar expression measurement.
9.4.2.4 Cluster Analysis
Cluster analysis (or simply “clustering”) has become a very popular method to
detect hidden structures in multivariate microarray data, and to detect co-regulated
genes. The popularity of cluster-analysis is understandable, as it requires no or few
prior hypotheses about the data. The application of cluster analysis is motivated by
the “guilt-by-association” assumption. If genes share a common mechanism of regulation (for example the same transcription factors) they could also be functionally
related; hence, it is useful to find groups of genes with similar expression profiles.
Clustering algorithms group measurements of gene expression into clusters. The
cluster assignment may be a hard assignment to a single class or a weighted gradual
assignment to multiple classes. The measurements are compared by their pair-wise
dissimilarity. Therefore, the notion of dissimilarity or distance measure plays a
central role for all clustering algorithms.
Hierarchical clustering algorithms construct a hierarchy of similar objects that
can be represented as a rooted binary tree dendrogram. Two types of hierarchical
clustering exist: agglomerative and divisive clustering. Agglomerative clustering is
the most popular approach.
Agglomerative clustering uses a bottom up approach. All objects start as singleton clusters and the most similar clusters are joined to form bigger clusters in
each step. The opposite approach is taken in divisive clustering algorithms where
all the genes initially constitute a single large cluster and are divided into smaller
and smaller clusters. The use of hierarchical clustering for the analysis of microarray data was popularized by Eisen et al. (1998) who also developed a software tool
to perform cluster analyses on microarrays. An appealing method for visualisation
of the results of the hierarchical clustering as a heat map is also presented in this
publication. The expression values are represented by colour codes; a red-green representation is used to denote the measured values. Negative log-ratios are projected
as green values and positive as red values, yielding black for values close to zero.
9.4.2.5 Classification
Sometimes, there is prior information about the origin of a certain sample. In such
cases, microarray data can be used to predict class information from the expression
profiles. The process of classification can be defined as assigning measurements to
discrete class labels.
All classification methods share the concept of a training phase and a classification phase. The data with known origin is used as the training set. Then a sample
with unknown origin can be classified into a class respective to its origin. During the
365
not be the case, for example, the Wilcoxon’s Rank Sum Statistic, or the Significance
Analysis of Microarrays (SAM) (Tusher et al. 2001) method. Other methods, such as
the CyberT method (Baldi and Long 2001), attempt to address the problem of small
numbers of replicates. The principal of CyberT is to borrow variance information
from other genes on the microarray having a similar expression measurement.
9.4.2.4 Cluster Analysis
Cluster analysis (or simply “clustering”) has become a very popular method to
detect hidden structures in multivariate microarray data, and to detect co-regulated
genes. The popularity of cluster-analysis is understandable, as it requires no or few
prior hypotheses about the data. The application of cluster analysis is motivated by
the “guilt-by-association” assumption. If genes share a common mechanism of regulation (for example the same transcription factors) they could also be functionally
related; hence, it is useful to find groups of genes with similar expression profiles.
Clustering algorithms group measurements of gene expression into clusters. The
cluster assignment may be a hard assignment to a single class or a weighted gradual
assignment to multiple classes. The measurements are compared by their pair-wise
dissimilarity. Therefore, the notion of dissimilarity or distance measure plays a
central role for all clustering algorithms.
Hierarchical clustering algorithms construct a hierarchy of similar objects that
can be represented as a rooted binary tree dendrogram. Two types of hierarchical
clustering exist: agglomerative and divisive clustering. Agglomerative clustering is
the most popular approach.
Agglomerative clustering uses a bottom up approach. All objects start as singleton clusters and the most similar clusters are joined to form bigger clusters in
each step. The opposite approach is taken in divisive clustering algorithms where
all the genes initially constitute a single large cluster and are divided into smaller
and smaller clusters. The use of hierarchical clustering for the analysis of microarray data was popularized by Eisen et al. (1998) who also developed a software tool
to perform cluster analyses on microarrays. An appealing method for visualisation
of the results of the hierarchical clustering as a heat map is also presented in this
publication. The expression values are represented by colour codes; a red-green representation is used to denote the measured values. Negative log-ratios are projected
as green values and positive as red values, yielding black for values close to zero.
9.4.2.5 Classification
Sometimes, there is prior information about the origin of a certain sample. In such
cases, microarray data can be used to predict class information from the expression
profiles. The process of classification can be defined as assigning measurements to
discrete class labels.
All classification methods share the concept of a training phase and a classification phase. The data with known origin is used as the training set. Then a sample
with unknown origin can be classified into a class respective to its origin. During the
