122
7.10 Data Import, Pre-Processing, and Visualization
Once IMS data are acquired, they are typically converted into imzML format, which
is an open source format allowing for analysis with a variety of software applications. Data can then be pre-processed using normalization, baseline correction, peak
detection, and peak alignment.
IMS data is high dimensional and large in terms of data size, often requiring a
form of data compression to decrease the computational load. Common methods to
compress data include binning mass spectra for each pixel or compressing based on
regions of interest (ROI). One compression method based on regions of interest is
image segmentation, which involves partitioning tissue into regions of homogenous
spectral profiles and identifying co-localized m/z values. Another data compression
approach uses unsupervised clustering to reduce dimensionality and extract features
for statistical analysis. A common unsupervised approach is principal component
analysis (PCA), which involves collapsing individual variables into groups that
exhibit similar behavior.
These groups are known as “components” and explain the variance in data with
far fewer dimensions. Because of its utility in dimensionality reduction, PCA is also
often used for downstream statistical analysis of pre-processed IMS data. Other
clustering approaches such as hierarchical clustering and k-means clustering can
also be used to visualize potential groupings of the data. Resultant data can be visualized by selecting individual m/z values, extracting the intensity of that m/z value
from each pixel’s spectrum, and plotting the intensity values as a heat map showing
the relative distribution of the m/z value across the tissue.
7.11 Statistical and Multivariate Analysis
Once the data have been processed and initial visualizations complete, a variety of
statistical techniques can be applied to compare data. If the experimental hypothesis
involves a direct comparison of relative m/z value intensities across regions of tissue, tests of significance such as t-tests and ANOVA can be applied. A t-test compares two groups of data, and ANOVA compares three or more. Unsupervised
approaches such as PCA can also be used to process the data. For instance, if the
regions of tissue are distinct in terms of their spectra, they should separate into components of the PCA, allowing distinct separations among the data. However, unsupervised approaches are usually applied to all m/z values identified in a tissue. If the
purpose of the study is to compare how well a specific m/z value can be used to
classify between groups or regions of tissue, a receiver operator characteristic curve
(ROC) analysis can be used. An ROC analysis is a test of accuracy and involves
plotting the true positive rate against the false positive rate. ROC analysis is often
used to determine if a molecular species with an explicit m/z is specific to a region.
J. C. McMillen et al.
Précédent

- 135/286

Suivant