Chapter 3
Basics of Neural Networks
Abstract In this chapter, we derive neural networks from the viewpoint of physical
models. A neural network is a nonlinear function that maps an input to an output,
and giving the network is equivalent to giving a function called an error function
in the case of supervised learning. By considering the output as dynamical degrees
of freedom and the input as an external field, various neural networks and their
deepened versions are born from simple Hamiltonians. Training (learning) is a
procedure for reducing the value of the error function, and we will learn the
specific method of backpropagation using the bra-ket notation popular in quantum
mechanics. And we will look at how the “universal approximation theorem” works,
which is why neural networks can express connections between various types of
data.
Now, let us move on to the explanation of supervised learning using neural networks.
Unlike many deep learning textbooks, this chapter describes machine learning in
terms of classical statistical physics. 1
3.1 Error Function from Statistical Mechanics
Error function is the log term of the error described by Kullback Leibler divergence
which we introduced in (2.8) and (2.9) in the previous chapter. The general
direction of explanation in many textbooks is to discuss the generalization error
and empirical error after defining the error function, but in this book, we will
explain that starting from (2.8) and (2.9), by considering the problem settings of
supervised learning as an appropriate statistical mechanical system, an appropriate
error function can be derived for each problem. In the process, we will see that
the structure of nonlinear functions such as sigmoid functions used in deep learning
1 For conventional explanations, we recommend Ref. [20]. Reading them together with this book
may complement understanding.
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2021
A. Tanaka et al., Deep Learning and Physics, Mathematical Physics Studies,
https://doi.org/10.1007/978-981-33-6108-9_3
35
Basics of Neural Networks
Abstract In this chapter, we derive neural networks from the viewpoint of physical
models. A neural network is a nonlinear function that maps an input to an output,
and giving the network is equivalent to giving a function called an error function
in the case of supervised learning. By considering the output as dynamical degrees
of freedom and the input as an external field, various neural networks and their
deepened versions are born from simple Hamiltonians. Training (learning) is a
procedure for reducing the value of the error function, and we will learn the
specific method of backpropagation using the bra-ket notation popular in quantum
mechanics. And we will look at how the “universal approximation theorem” works,
which is why neural networks can express connections between various types of
data.
Now, let us move on to the explanation of supervised learning using neural networks.
Unlike many deep learning textbooks, this chapter describes machine learning in
terms of classical statistical physics. 1
3.1 Error Function from Statistical Mechanics
Error function is the log term of the error described by Kullback Leibler divergence
which we introduced in (2.8) and (2.9) in the previous chapter. The general
direction of explanation in many textbooks is to discuss the generalization error
and empirical error after defining the error function, but in this book, we will
explain that starting from (2.8) and (2.9), by considering the problem settings of
supervised learning as an appropriate statistical mechanical system, an appropriate
error function can be derived for each problem. In the process, we will see that
the structure of nonlinear functions such as sigmoid functions used in deep learning
1 For conventional explanations, we recommend Ref. [20]. Reading them together with this book
may complement understanding.
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2021
A. Tanaka et al., Deep Learning and Physics, Mathematical Physics Studies,
https://doi.org/10.1007/978-981-33-6108-9_3
35
