Research on UAV Cluster’s Operation
Strategy Based on Reinforcement
Learning Approach
Yi Mao
(&) and Yuxin Hu
State Key Laboratory of Air Traffic Management System and Technology,
Nanjing 210007, China
mao_y@nuaa.edu.cn
Abstract. It is of necessity to formulate an overall UAV regulation scheme that
covers each UAV’s path and task implementation in the operation of UAV
cluster. However, failure to fully realize consistency with pre-planned scheme in
actual task implementation may occur considering changes of task, damage,
addition or reduction of UAVs, fuel loss, unknown circumstance and other
uncertainties, which thus entail a simultaneous online regulation scheme. By
predicting UAV’s 4D track, posture, task and demand of resources in regulation
and clarifying data set in UAV cluster operation, quick planning of UAV’s flight
path and operation can be realized, thus reducing probability of scheme
adjustment and improving operation efficiency.
Keywords: UAV Á Operation strategy Á Reinforcement learning approach
1 Introduction
UAV cluster operation is a complex multitask system that is affected and restricted by
UAV’s functional performance, load, information context, support equipment and other
factors. Dynamic regulation of UAV cluster operation involves strategic training and
optimization at the stage of gathering in takeoff, formation in flight, transformation of
flight form, dispersion and collection of multiple UAVs of different types and structures
[1–5].
Traditional training approaches, in general, include linear planning, dynamic
planning, branch-and-bound method, elimination method and other conventional
training optimization approaches frequently adopted in operational researches. In the
application of optimal approach, the problem is often simplified for the sake of
mathematical description and modeling so as to formulate an optimized regulation
scheme [6–9]. UAV’s specification, load, ammunition, coverage of support resources
and battlefield context are all critical factors that impose influences in UAV cluster
operation. Coupled with randomness and dynamicity in operation and a number of reregulation tasks, UAV cluster operation is a NP-hard problem, characterized with great
difficulties in solution, long time spent on solution and inability to realize simultaneous
online UAV strategic cluster training [10–13].
© Springer Nature Singapore Pte Ltd. 2020
Q. Liang et al. (Eds.): Artificial Intelligence in China, LNEE 572, pp. 28–35, 2020.
https://doi.org/10.1007/978-981-15-0187-6_4
Précédent

- 40/679

Suivant