28
2 Introduction to Machine Learning
Column: Probability Theory and Information Theory
As explained in this book, probability theory and information theory are indispensable for developing the theory of machine learning. This column explains some
important concepts.
Joint and Conditional Probabilities
Machine learning methods often consider the probability of the input value x and
the teaching signal d. The joint probability represents the appearance probability
of these two variables. Let us write it as
P (x, d) .
(2.25)
Of course, we have
1 =
x
d
P (x, d) .
(2.26)
Here,
is a sum when x, d are discrete valued, and is replaced by an integral when
they are continuous variables.
There should be some situations when we do not want to look at the value of x
but want to consider the probability of d. It can be expressed as
P (d) =
x
P (x, d).
(2.27)
The summation operation is called marginalizing, and the probability P (d) is called
marginal probability.
In some cases, d may already be given. This situation is represented by
P (x|d): probability of x when d is given.
(2.28)
This is called conditional probability. In this case, d is regarded as a fixed quantity,
1 =
x
P (x|d) .
(2.29)
So, in fact, we find
P (x|d) =
P (x, d)
P (d)
=
P (x, d)
X P (X, d)
.
(2.30)
Précédent

- 38/211

Suivant