72
biological processes must be differentiated from extraneous ones. For example,
Mahieu et al. showed that although a dataset had over 25,000 features, with the
subtraction of isotopes, adducts, artifacts, and contaminants, less than 1000 were
metabolites [70]. Preprocessing methods facilitate recognition of data as either
meaningful or irrelevant.
There are a variety of tools for data preprocessing such as noise filtering, spectral
deconvolution, chromatogram alignment, and retention time correction. Data processing such as peak detection, peak alignment, metabolite identification, quality
control, normalization, statistical analysis, metabolite quantification, and in silico
fragmentation are also used [71]. Both targeted and untargeted metabolomics methods share similar data preprocessing. With targeted methods such as MRM, chromatographic features linked to specific MS/MS transitions are often used. A variety
of commercial software, such as LCQuan (ThermoFisher Scientific), allow for the
identification of internal standards for relative quantitation or the import of calibration curves for absolute quantitation [72].
To reduce the dimensionality of the data, classification and clustering tools are
used [73]. Once metabolites are identified, relative or absolute quantitation can be
performed to determine the overall role of observed metabolic changes in a global
framework. There are a variety of commercially available tools for targeted and
untargeted LC-MS data analysis including LCQuan, Agilent Masshunter, Bruker’s
Profile Analysis, Thermo SIEVE, Waters’ Progensis QI and more [72, 74–77]. In
addition there are a number of open source, vendor-independent tools including
XCMS/XCMS Online, Mzmine 2, and MS-DIAL [78–80].
After data preprocessing, metabolite identification remains challenging owing to
incomplete spectral libraries and incompatibility between databases and data types.
For example, some databases are compatible with MS
n
data while others are only
designed to search compounds. Table 4.1 provides a list of relevant spectral libraries
and databases to assist in the identification of metabolites. When determining which
database best fits an experiment, it is important not only to consider the total number
of compounds and the data type, but also the original data used to build the data
base. For example, it is possible to limit false identifications for a human-based
experiment by selecting HMDB rather than an in silco prediction-based database.
Once metabolites are identified, pathways become integral in identifying the collective role metabolites play in relation to a scientific question. Table 4.2 lists a few
metabolic pathway analysis tools, their number of reference pathways and the number of organisms on which they are based. Using these tools, data sets can be mapped
into known biological networks to aid in the interpretation of the results and provide
context that will help generate future hypotheses. Although pathway analysis can
often bring about answers to a variety of biological questions, it is important to note
that many experiments are temporal and looking at the accumulation of metabolites
or the change between experimental groups. In order to directly follow a metabolic
pathway, heavy labeling experiments are needed [87].
E. S. Rivera et al.
biological processes must be differentiated from extraneous ones. For example,
Mahieu et al. showed that although a dataset had over 25,000 features, with the
subtraction of isotopes, adducts, artifacts, and contaminants, less than 1000 were
metabolites [70]. Preprocessing methods facilitate recognition of data as either
meaningful or irrelevant.
There are a variety of tools for data preprocessing such as noise filtering, spectral
deconvolution, chromatogram alignment, and retention time correction. Data processing such as peak detection, peak alignment, metabolite identification, quality
control, normalization, statistical analysis, metabolite quantification, and in silico
fragmentation are also used [71]. Both targeted and untargeted metabolomics methods share similar data preprocessing. With targeted methods such as MRM, chromatographic features linked to specific MS/MS transitions are often used. A variety
of commercial software, such as LCQuan (ThermoFisher Scientific), allow for the
identification of internal standards for relative quantitation or the import of calibration curves for absolute quantitation [72].
To reduce the dimensionality of the data, classification and clustering tools are
used [73]. Once metabolites are identified, relative or absolute quantitation can be
performed to determine the overall role of observed metabolic changes in a global
framework. There are a variety of commercially available tools for targeted and
untargeted LC-MS data analysis including LCQuan, Agilent Masshunter, Bruker’s
Profile Analysis, Thermo SIEVE, Waters’ Progensis QI and more [72, 74–77]. In
addition there are a number of open source, vendor-independent tools including
XCMS/XCMS Online, Mzmine 2, and MS-DIAL [78–80].
After data preprocessing, metabolite identification remains challenging owing to
incomplete spectral libraries and incompatibility between databases and data types.
For example, some databases are compatible with MS
n
data while others are only
designed to search compounds. Table 4.1 provides a list of relevant spectral libraries
and databases to assist in the identification of metabolites. When determining which
database best fits an experiment, it is important not only to consider the total number
of compounds and the data type, but also the original data used to build the data
base. For example, it is possible to limit false identifications for a human-based
experiment by selecting HMDB rather than an in silco prediction-based database.
Once metabolites are identified, pathways become integral in identifying the collective role metabolites play in relation to a scientific question. Table 4.2 lists a few
metabolic pathway analysis tools, their number of reference pathways and the number of organisms on which they are based. Using these tools, data sets can be mapped
into known biological networks to aid in the interpretation of the results and provide
context that will help generate future hypotheses. Although pathway analysis can
often bring about answers to a variety of biological questions, it is important to note
that many experiments are temporal and looking at the accumulation of metabolites
or the change between experimental groups. In order to directly follow a metabolic
pathway, heavy labeling experiments are needed [87].
E. S. Rivera et al.
