5 Conclusion
At present, there are mainly two types of reinforcement learning algorithms to solve
UAV cluster’s independent intelligent operation problems: first, value function estimation, which is adopted in reinforcement learning researches mostly extensively with
quickest development; second, direct strategy space search approach, such as genetic
algorithm, genetic program design, simulated annealing and other evolution
approaches.
Direct strategy space search approach and value function estimation approach can
be both used to solve reinforcement learning problems. Both improve Agent strategy
by means of training learning. However, direct strategy space search approach does not
use strategy as a mapping function from state to movement. Value function is not taken
into consideration. In another word, environment status is not taken into consideration.
Value function estimation approach concentrates on value function; that is, environment status is regarded as a core element. Value function estimation approach can
realize learning based on the interaction between Agent and environment, regardless of
quality of strategy—either good or bad. On the contrary, direct strategy space search
approach fails to realize such a segmented gradual learning. It is effective for reinforcement learning with enough small strategy space or sound structure that facilitates
easiness to find out an optimal strategy. In addition, when Agent fails to precisely
perceive environment status, direct strategy space search approach displays great
advantages. By modeling UAV cluster’s independent intelligent operation, we can find
that value function estimation is better for large-scale complex problems because it can
make better use of effective computing resources and reach the goal of solving UAV
cluster’s independent intelligent operation.
References
1. Raivio T (2001) Capture set computation of an optimally guided missile [J]. J Guid Control
Dyn 24(6):1167–1175
2. Fei A (2017) Analysis on related issues about resilient command and control system design.
Command Inf Syst Technol 8(2):1–4
3. Smart WD (2004) Explicit manifold representations for value-function approximation in
reinforcement learning. In: AMAI
4. Keller PW, Mannor S, Precup D (2006) Automatic basis function construction for
approximate dynamic programming and reinforcement learning. In: Proceedings of the 23rd
international conference on Machine learning. ACM, pp 449–456
5. Dearden R, Friedman N, Russell S (1998) Bayesian Q-learning. In: AAAI/IAAI, pp 761–768
6. Chase HW, Kumar P, Eickhoff SB et al (2015) Reinforcement learning models and their
neural correlates: an activation likelihood estimation meta-analysis. Cogn Affect Behav
Neurosci 15(2):435–459
7. Botvinick M (2012) Hierarchical reinforcement learning and decision making. Curr Opin
Neurobiol 22(6):956–962
8. Smith RE, Dike BA, Mehra RK et al (2000) Classifier systems in combat: two-sided learning
of maneuvers for advanced fighter aircraft. Comput Methods Appl Mech Eng 186(2):421–
437
34
Y. Mao and Y. Hu
At present, there are mainly two types of reinforcement learning algorithms to solve
UAV cluster’s independent intelligent operation problems: first, value function estimation, which is adopted in reinforcement learning researches mostly extensively with
quickest development; second, direct strategy space search approach, such as genetic
algorithm, genetic program design, simulated annealing and other evolution
approaches.
Direct strategy space search approach and value function estimation approach can
be both used to solve reinforcement learning problems. Both improve Agent strategy
by means of training learning. However, direct strategy space search approach does not
use strategy as a mapping function from state to movement. Value function is not taken
into consideration. In another word, environment status is not taken into consideration.
Value function estimation approach concentrates on value function; that is, environment status is regarded as a core element. Value function estimation approach can
realize learning based on the interaction between Agent and environment, regardless of
quality of strategy—either good or bad. On the contrary, direct strategy space search
approach fails to realize such a segmented gradual learning. It is effective for reinforcement learning with enough small strategy space or sound structure that facilitates
easiness to find out an optimal strategy. In addition, when Agent fails to precisely
perceive environment status, direct strategy space search approach displays great
advantages. By modeling UAV cluster’s independent intelligent operation, we can find
that value function estimation is better for large-scale complex problems because it can
make better use of effective computing resources and reach the goal of solving UAV
cluster’s independent intelligent operation.
References
1. Raivio T (2001) Capture set computation of an optimally guided missile [J]. J Guid Control
Dyn 24(6):1167–1175
2. Fei A (2017) Analysis on related issues about resilient command and control system design.
Command Inf Syst Technol 8(2):1–4
3. Smart WD (2004) Explicit manifold representations for value-function approximation in
reinforcement learning. In: AMAI
4. Keller PW, Mannor S, Precup D (2006) Automatic basis function construction for
approximate dynamic programming and reinforcement learning. In: Proceedings of the 23rd
international conference on Machine learning. ACM, pp 449–456
5. Dearden R, Friedman N, Russell S (1998) Bayesian Q-learning. In: AAAI/IAAI, pp 761–768
6. Chase HW, Kumar P, Eickhoff SB et al (2015) Reinforcement learning models and their
neural correlates: an activation likelihood estimation meta-analysis. Cogn Affect Behav
Neurosci 15(2):435–459
7. Botvinick M (2012) Hierarchical reinforcement learning and decision making. Curr Opin
Neurobiol 22(6):956–962
8. Smith RE, Dike BA, Mehra RK et al (2000) Classifier systems in combat: two-sided learning
of maneuvers for advanced fighter aircraft. Comput Methods Appl Mech Eng 186(2):421–
437
34
Y. Mao and Y. Hu
