71
Table 4.3 Preprocessing steps of a lipidomic data analysis
Preprocessing steps
Description
Peak detection
It involves the detection of each lipid metabolite peak
from the raw data for each individual sample. A list of
peaks characterized by its mass (m/z) and retention time
is obtained. Subsequently, peak quantification is done
by determining the peak area through peak integration,
and a detection limit (signal-to-noise level) for the
intensity is set a priori to eliminate noise peaks in the
case of LC-MS-based lipidomic data. For direct
infusion data, the intensity of each peak is determined
as a (weighted) average over the collection time
Peak quantification/normalization
It is accomplished using internal standards in order to
obtain semiquantitative data that can be further
statistically analyzed. Multiple internal standards are
used to counteract the matrix effects and suppression in
response of different lipid classes
Peak grouping
Lipid metabolite peaks are grouped according to their
peak positions across groups of samples for their
statistical comparison. Peak matching procedures search
for peaks across samples within a group of related
samples falling within a pre-specified m/z and retention
time distance of each other and, subsequently, represent
the same metabolite. Nonlinear time alignment
approaches are further applied to correct time shifts
between samples
Imputation of missing peaks
One or few samples in a group of related samples may
sometimes miss one or more metabolite peaks due to
sample outliers or experimental/bioinformatic
shortcomings. Such missing peaks can be added to the
peak table by assuming that they are located at the same
position as the identified peaks in the related samples
and integrate the measured intensity within these areas.
At this stage, a peak group in the list is referred to as a
feature, defined by its m/z and RT and intensity. Further
analysis of the dataset determines which features
correspond to identified metabolites or noise and which
remain unidentified
Visualization
The complexity of lipidomic data decrees the use of
visualization methods to explore raw and processed
data for quality control and validation of results. For
every feature, simple graphs of the contributing peaks
(overlays of extracted ion chromatograms) or box plots
of the intensities are employed to obtain rapid insight
into the shape of the peaks and the relative differences
between groups of samples
Isotope correction
At low resolution, often peaks from isotope molecules
of one lipid overlap with peaks of a lipid with two more
hydrogen atoms. Thus, the intensity of the isotope peak
needs to be subtracted from the intensity of the second
lipid, before all lipids can be properly quantified
(continued)
4 Seaweed Lipidomics in the Era of ‘Omics’ Biology: A Contemporary Perspective
Table 4.3 Preprocessing steps of a lipidomic data analysis
Preprocessing steps
Description
Peak detection
It involves the detection of each lipid metabolite peak
from the raw data for each individual sample. A list of
peaks characterized by its mass (m/z) and retention time
is obtained. Subsequently, peak quantification is done
by determining the peak area through peak integration,
and a detection limit (signal-to-noise level) for the
intensity is set a priori to eliminate noise peaks in the
case of LC-MS-based lipidomic data. For direct
infusion data, the intensity of each peak is determined
as a (weighted) average over the collection time
Peak quantification/normalization
It is accomplished using internal standards in order to
obtain semiquantitative data that can be further
statistically analyzed. Multiple internal standards are
used to counteract the matrix effects and suppression in
response of different lipid classes
Peak grouping
Lipid metabolite peaks are grouped according to their
peak positions across groups of samples for their
statistical comparison. Peak matching procedures search
for peaks across samples within a group of related
samples falling within a pre-specified m/z and retention
time distance of each other and, subsequently, represent
the same metabolite. Nonlinear time alignment
approaches are further applied to correct time shifts
between samples
Imputation of missing peaks
One or few samples in a group of related samples may
sometimes miss one or more metabolite peaks due to
sample outliers or experimental/bioinformatic
shortcomings. Such missing peaks can be added to the
peak table by assuming that they are located at the same
position as the identified peaks in the related samples
and integrate the measured intensity within these areas.
At this stage, a peak group in the list is referred to as a
feature, defined by its m/z and RT and intensity. Further
analysis of the dataset determines which features
correspond to identified metabolites or noise and which
remain unidentified
Visualization
The complexity of lipidomic data decrees the use of
visualization methods to explore raw and processed
data for quality control and validation of results. For
every feature, simple graphs of the contributing peaks
(overlays of extracted ion chromatograms) or box plots
of the intensities are employed to obtain rapid insight
into the shape of the peaks and the relative differences
between groups of samples
Isotope correction
At low resolution, often peaks from isotope molecules
of one lipid overlap with peaks of a lipid with two more
hydrogen atoms. Thus, the intensity of the isotope peak
needs to be subtracted from the intensity of the second
lipid, before all lipids can be properly quantified
(continued)
4 Seaweed Lipidomics in the Era of ‘Omics’ Biology: A Contemporary Perspective
