192 Beam-based Correction and Optimization for Accelerators
improvement) [86], or by maximizing the expected amount of improvement
(EI, expected improvement) [82], or by minimizing the lower confidence bound
(LCB) [4]. The GP-LCB acquisition function is defined as
GP-LCB(x) = µ(x) − κ t σ(x),
(7.29)
where the value of κ t is chosen to balance between the the exploration strategy
(with a large κ t ) and the exploitation strategy (with a small κ t ). After evaluation, the new trial solution enters the sample data set, D t , which is then used
to update the model. As more data are collected, the approximation of the
objective function becomes more accurate, which would lead the algorithm to
converge to the optimum.
Evaluation of the acquisition function requires the inversion of the kernel
matrix. Subsequent calculations with the Gaussian process model involve matrix multiplications. As the number of data points increases, the computation
time will also increase. This is acceptable if the number of data points is on
the order of hundreds. But the computation would become too slow if the
data sample size is much larger.
When noise is present in the function evaluations, the diagonal elements
of the kernel matrix change as the random noise enters the variance, but the
off-diagonal elements do not change (as the noise at different sample point is
independent), hence in Eq. (7.28) the kernel matrix is replaced with,
K → K + σ
2 I.
(7.30)
Noise affects the accuracy and even the validity of the Gaussian process model
and in turn the performance of the optimizer.
Multi-generation Gaussian process optimizer (MG-GPO): The
ability of predicting function values with the posterior GP model can substantially enhance the efficiency of the stochastic optimization algorithms.
The MG-GPO [56] algorithm is a method that utilizes this ability to select
trial solutions with high potential for the actual function evaluation. It operates iteratively and maintains a fixed number of good solutions, in the same
manner as the GA and PSO algorithms. In the case of a multi-objective optimization, a GP model is built for each objective function. Many new solutions
are generated for each good solution using cross-over and mutation operations.
These solutions are tested with the posterior GP models and are ranked with
non-dominated sorting according to the corresponding acquisition function
values. Only a fixed number of solutions in the leading fronts are evaluated on
the real system. The evaluated solutions are then combined with the existing
good solutions, from which a new population of good solutions are selected
for the next generation.
The GP models are rebuilt at each generation using only the good solution
population and the recently evaluated solutions. Therefore, the size of the
sample data set does not grow indefinitely. In simulation, it was shown that
the GP-GPO method outperforms both the NSGA-II and MOPSO methods.
improvement) [86], or by maximizing the expected amount of improvement
(EI, expected improvement) [82], or by minimizing the lower confidence bound
(LCB) [4]. The GP-LCB acquisition function is defined as
GP-LCB(x) = µ(x) − κ t σ(x),
(7.29)
where the value of κ t is chosen to balance between the the exploration strategy
(with a large κ t ) and the exploitation strategy (with a small κ t ). After evaluation, the new trial solution enters the sample data set, D t , which is then used
to update the model. As more data are collected, the approximation of the
objective function becomes more accurate, which would lead the algorithm to
converge to the optimum.
Evaluation of the acquisition function requires the inversion of the kernel
matrix. Subsequent calculations with the Gaussian process model involve matrix multiplications. As the number of data points increases, the computation
time will also increase. This is acceptable if the number of data points is on
the order of hundreds. But the computation would become too slow if the
data sample size is much larger.
When noise is present in the function evaluations, the diagonal elements
of the kernel matrix change as the random noise enters the variance, but the
off-diagonal elements do not change (as the noise at different sample point is
independent), hence in Eq. (7.28) the kernel matrix is replaced with,
K → K + σ
2 I.
(7.30)
Noise affects the accuracy and even the validity of the Gaussian process model
and in turn the performance of the optimizer.
Multi-generation Gaussian process optimizer (MG-GPO): The
ability of predicting function values with the posterior GP model can substantially enhance the efficiency of the stochastic optimization algorithms.
The MG-GPO [56] algorithm is a method that utilizes this ability to select
trial solutions with high potential for the actual function evaluation. It operates iteratively and maintains a fixed number of good solutions, in the same
manner as the GA and PSO algorithms. In the case of a multi-objective optimization, a GP model is built for each objective function. Many new solutions
are generated for each good solution using cross-over and mutation operations.
These solutions are tested with the posterior GP models and are ranked with
non-dominated sorting according to the corresponding acquisition function
values. Only a fixed number of solutions in the leading fronts are evaluated on
the real system. The evaluated solutions are then combined with the existing
good solutions, from which a new population of good solutions are selected
for the next generation.
The GP models are rebuilt at each generation using only the good solution
population and the recently evaluated solutions. Therefore, the size of the
sample data set does not grow indefinitely. In simulation, it was shown that
the GP-GPO method outperforms both the NSGA-II and MOPSO methods.
