256
C. Karul . S. Soyupak
always be reduced at each iteration of the algorithm. An initial ~ value of 0.001
was used.
13.3.1.2
Data Pre-Processing
To increase the efficiency of training, the network inputs and targets were scaled
by normalizing the mean and S.D. of the training set. This process normalizes the
input and target values so that they have zero mean and unity S.D. When training
is completed, the simulation results are de-normalized by reversing the action.
13.3.1.3
Improving Generalization
A three-Iayer feed-forward backpropagation neural network with sufficient
number of neurons can approximate any function. Thus, one should be aware of
the danger that neural network may be memorizing the available data rather than
generalizing it, so called over-fitting the data. An over-fitted neural network
model typically imitates the data in the training set very successfully but generates
a bad estimation for the data not included in the training. For a good
generalization, over-fitting should be prevented taking the appropriate measures.
Over-fitting can be prevented by utilizing either of the two methods: i)
Regularization, and ii) Early stopping. The second method, early stopping, is used
in this study to prevent overtraining.
To decide when to stop training, the data set is randomly divided into three subsets, one half is used for training, one quarter for validation and the last quarter for
testing. The error term, i.e. the difference between measured target values and the
calculated values was calculated for the training set, validation set and the test set
separately. The error on the validation set will normally decrease during the initial
part of the training. However, when the network begins to overfit the data, the
validation set error will start to rise. When this increase continues for a predefined
number of iterations the training is stopped and the weight values are kept
constant. The set is used to compare with the validation set to see if they exhibit a
similar behaviour. If the validation errors and test errors do not show a similar
behaviour, this may indicate a poor division of data.
Mean square error is the typical performance function used in feed-forward
neural networks:
(13.6)
where MSE is the mean square error; N is the number of elements; i is the index
for elements; ej is the error of the i
th element; tj is the target value (measured) for
i
th element and a is the calculated value for i th element.
•
I
C. Karul . S. Soyupak
always be reduced at each iteration of the algorithm. An initial ~ value of 0.001
was used.
13.3.1.2
Data Pre-Processing
To increase the efficiency of training, the network inputs and targets were scaled
by normalizing the mean and S.D. of the training set. This process normalizes the
input and target values so that they have zero mean and unity S.D. When training
is completed, the simulation results are de-normalized by reversing the action.
13.3.1.3
Improving Generalization
A three-Iayer feed-forward backpropagation neural network with sufficient
number of neurons can approximate any function. Thus, one should be aware of
the danger that neural network may be memorizing the available data rather than
generalizing it, so called over-fitting the data. An over-fitted neural network
model typically imitates the data in the training set very successfully but generates
a bad estimation for the data not included in the training. For a good
generalization, over-fitting should be prevented taking the appropriate measures.
Over-fitting can be prevented by utilizing either of the two methods: i)
Regularization, and ii) Early stopping. The second method, early stopping, is used
in this study to prevent overtraining.
To decide when to stop training, the data set is randomly divided into three subsets, one half is used for training, one quarter for validation and the last quarter for
testing. The error term, i.e. the difference between measured target values and the
calculated values was calculated for the training set, validation set and the test set
separately. The error on the validation set will normally decrease during the initial
part of the training. However, when the network begins to overfit the data, the
validation set error will start to rise. When this increase continues for a predefined
number of iterations the training is stopped and the weight values are kept
constant. The set is used to compare with the validation set to see if they exhibit a
similar behaviour. If the validation errors and test errors do not show a similar
behaviour, this may indicate a poor division of data.
Mean square error is the typical performance function used in feed-forward
neural networks:
(13.6)
where MSE is the mean square error; N is the number of elements; i is the index
for elements; ej is the error of the i
th element; tj is the target value (measured) for
i
th element and a is the calculated value for i th element.
•
I
