Chapter 2
Introduction to Machine Learning
Abstract In this chapter, we learn the general theory of machine learning. We shall
take a look at examples of what learning is, what is the meaning of “machines
learned,” and what relative entropy is. We will learn how to handle data in
probability theory, and describe “generalization” and its importance in learning.
2.1 The Purpose of Machine Learning
Deep learning is a branch of what is called machine learning. The old “definition”
of machine learning by computer scientist Arthur Samuel [16] is “Field of study
that gives computers the ability to learn without being explicitly programmed.” In
more modern terms, “From experience alone, let the machine automatically gain the
ability to extract the structure behind it and apply it to unknown situations.” 1
There are various methods depending on how the experience is given to the
machine, so here let us assume a situation in which the experience is already stored
in a database. In this context, there are two types of data to consider:
• Supervised data{(x[i], d[i])} i=1,2,...,#
• Unsupervised data{(x[i], − )} i=1,2,...,#
Here x[i] denotes the i-th data value (generally vector valued), and d[i] denotes
teaching signal. The mark “−” indicates that no teaching signal is given, and in the
case of unsupervised data, only {x[i]} is actually provided. # indicates the number
of data. See, for example, Fig. 2.1.
1 Richard Feynman said [17], “We can imagine that this complicated array of moving things which
consists “the world” is something like a great chess game being played by the gods, and we are
observers of the game. We do not know what the rules of the game are; all we are allowed to do
is to watch the playing. Of course, if we watch long enough, we may eventually catch on to a few
of the rules.” It is a good parable that captures the essence of the inverse problem of guessing rules
and structures. The machine learning, compared with this example, would mean that the observer
is a machine instead of a human.
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2021
A. Tanaka et al., Deep Learning and Physics, Mathematical Physics Studies,
https://doi.org/10.1007/978-981-33-6108-9_2
17
Introduction to Machine Learning
Abstract In this chapter, we learn the general theory of machine learning. We shall
take a look at examples of what learning is, what is the meaning of “machines
learned,” and what relative entropy is. We will learn how to handle data in
probability theory, and describe “generalization” and its importance in learning.
2.1 The Purpose of Machine Learning
Deep learning is a branch of what is called machine learning. The old “definition”
of machine learning by computer scientist Arthur Samuel [16] is “Field of study
that gives computers the ability to learn without being explicitly programmed.” In
more modern terms, “From experience alone, let the machine automatically gain the
ability to extract the structure behind it and apply it to unknown situations.” 1
There are various methods depending on how the experience is given to the
machine, so here let us assume a situation in which the experience is already stored
in a database. In this context, there are two types of data to consider:
• Supervised data{(x[i], d[i])} i=1,2,...,#
• Unsupervised data{(x[i], − )} i=1,2,...,#
Here x[i] denotes the i-th data value (generally vector valued), and d[i] denotes
teaching signal. The mark “−” indicates that no teaching signal is given, and in the
case of unsupervised data, only {x[i]} is actually provided. # indicates the number
of data. See, for example, Fig. 2.1.
1 Richard Feynman said [17], “We can imagine that this complicated array of moving things which
consists “the world” is something like a great chess game being played by the gods, and we are
observers of the game. We do not know what the rules of the game are; all we are allowed to do
is to watch the playing. Of course, if we watch long enough, we may eventually catch on to a few
of the rules.” It is a good parable that captures the essence of the inverse problem of guessing rules
and structures. The machine learning, compared with this example, would mean that the observer
is a machine instead of a human.
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2021
A. Tanaka et al., Deep Learning and Physics, Mathematical Physics Studies,
https://doi.org/10.1007/978-981-33-6108-9_2
17
