2.4.2 Machine Learning
The final strand of QCTFF to discuss briefly is the way atomic properties (both
intra- and inter-) are predicted from the coordinates of the nuclei surrounding the
atom of interest. In general, this mapping is so complex that it needs a machine
learning method, and one that can handle high-dimensional spaces, given the large
number of coordinates that influence the atom of interest. Kriging [78] is such a
method. Originating in geostatistics, Kriging is a powerful interpolation technique
that can capture the behaviour of an output as a function of many inputs, using a
relatively small amount of data points. In its infancy [79] it succeeded in predicting
where the best location for a mine would be in a two-dimensional landscape, based
on measurements of a precious material (originally gold but could be diamond, oil,
uranium or any ore) at various locations in this landscape. The basic idea of Kriging
is to predict the value of a function at a given point by computing a weighted
average of the known values of the function in the neighborhood of the point. An
accessible account of the details of Kriging as used within the QCTFF context has
been given elsewhere [80].
Here we highlight one key idea, namely that of maximising the likelihood L,
which has not been clarified in that previous account [80]. To fix thoughts, let us
start with a simple example: a coin is being tossed thrice. If the coin is fair, then the
probability to observe head up, (denoted p H ), is one half, that is p H = 0.5. Equally,
the probably of observing tail denoted p T is one half, or p T = 0.5. The probability to
observe head up twice and then tail (HHT) is p HHT = p H p H p T = p H
2 (1-p H ) = 0.125.
An equivalent way of saying this is to reverse this statement: the likelihood L that
the coin was fair (i.e. p H = 0.5), given the observation of two heads being up
(HHT), is one eighth, i.e. L = 0.125. This is formally written as follows:
Lðp H ¼ 0:5jHHTÞ ¼ 0:125
ð2:20Þ
In summary, the likelihood L is a function returning the probability of observed
outcomes (e.g. HHT), given a parameter value (i.e. p H ). We now ask ourselves how
the likelihood L = p H
2 (1-p H ) can be maximised. Mathematically this is easy: calculus
tells us that dL/dp H = d/dp H [p H
2 (1-p H )] = 2p H -3 p H
2 , which vanishes when p H = 2/3.
A plot, or a quick calculation of the second derivative, tells us that p H = 2/3 is
indeed a maximum, at which point L = 4/27. The result that p H, max L = 2/3 can be
intuitively understood by stating that the coin is biased towards heads up, by a
factor 2 over tail up. Indeed, with such a bias, the probability of the observed
outcomes HHT, given p H = 2/3, is maximal. How does all this help understanding a
key aspect behind Kriging?
Kriging uses the same strategy of maximising the likelihood: it finds the
parameters θ h and p h (h = 1, 2, …, d) in the so-called Gram matrix R,
2 On Quantum Chemical Topology
43
Précédent

- 51/582

Suivant