record, which we can explore to learn about the mutational processes that a tumor has been
exposed to (Fig. 2.2) [8].
Review Question 4
Which cancer types do you expect to have the highest and the lowest numbers of
mutations, and why?
Recently, computational methodologies have been developed to extract all these mutational signatures from any given mutation catalog C, which is a matrix that has samples as
columns and mutation classes as rows. The latter are the six possible mutation types (C:G
> A:T, C:G > G:C, C:G > T:A, T:A > A:T, T:A > C:G, and T:A > G:C) taken into their
trinucleotide contexts (i.e., the base before and after each mutation), thus yielding
6 Â 4 Â 4 ¼ 96 different mutation classes (Fig. 2.2). These algorithms try to optimally
solve C % SE, where S is the signature matrix (with mutation classes as rows and signatures
as columns) and E is the exposure matrix (with samples as columns and signatures as rows)
[9]. This way, we can learn which samples have which signatures with which “intensity”
(the exposure).
Mutational signature analysis has been tremendously useful in recent years to elucidate
the mechanism of action of several carcinogens (Fig. 2.3). For example, this type of
analysis revealed in 2013 the types of mutations induced by aristolochic acid, a substance
present in plants traditionally used for medicinal purposes, which are dominated by A>T
transversions by formation of aristolactam-DNA adducts [10]. Another example is the
discovery of the extent to which APOBEC enzymes play a role in cancer and associated
cell line models by their strong activity as cytidine deaminases [11]. However, this exciting
field is in constant evolution and several distinct bioinformatic methodologies have been
published for the extraction of mutational signatures from cancer genomes. The original
method, a de novo extraction algorithm based on non-negative matrix factorization
(NNMF), was published by Alexandrov and collaborators and applied to 7042 distinct
tumors from 30 different types of cancer, being able to identify 21 mutational signatures
[8]. Since then, other methods, both de novo and approaches fitting C to a matrix of known
signatures, have been published with varying results. Researchers doing this type of
analysis can run into problems such as ambiguous signature assignment, the fact that
some localized mutational processes may not be taken into account, and the algorithms’
assumption that all samples being analyzed have a similar mutational profile [9]. A recent
study has suggested that a combination of de novo and fitting approaches may reduce false
positives while still allowing the discovery of novel signatures [9]. Generally, as is the case
with any bioinformatic tool, researchers must be cautious about their results and should
perform a manual curation where possible, making use of prior biological knowledge and
making sure results make sense.
Mutational signature analysis has recently been expanded to include multiple-base
mutations and small indels. To date, nearly 24,000 cancer genomes and exomes have
been analyzed, with 49 single-base substitution (SBS), 11 doublet-base substitutions, four
clustered base substitutions, and 17 indel mutational signatures discovered [8].
2 Opportunities and Perspectives of NGS Applications in Cancer Research
23
Précédent

- 34/225

Suivant