3 Optimization of Operation Strategic Training Modeling
One key assumption is that UAV cluster’s operation strategy is based on the interaction
between UAV and environment. Such an interaction can be regarded as MDP (Markov
decision process). Optimization of operation strategy based on reinforcement learning
can adopt solutions to Markov problems. Supposing that the system is observed at time
t ¼ t 1; t 2; . . .; t n , one operation strategy’s decision process is composed of five elements:
S; A S
ð Þ; P
a
ss ; R
a
ss ; V S; S
0
2 S; a 2 A S
ð Þ
j
ð3:1Þ
Each element’s meaning is as below:
① S is the non-empty set composed of all possible states of the system. Sometimes, it
is also called “system’s state space,” which can be a definite, denumerable or
arbitrary non-empty set. S, S′ are elements of S, implying the states:
② s 2 S, A(s) is the set of all possible movements at states.
③ When system is at the states at decision-making time t, after executing decision a,
the probability of system’s S′ state at the next decision-making point t + 1 is P
a
ss .
The transition probability of all movements constitutes one transition matrix
cluster:
P
a
ss ¼ Pr S t þ 1 ¼ s
0
jS t ¼ S; a t ¼ a
f
g
ð3:2Þ
④ When system is at the states at decision-making time t, after executing decision a,
the immediate return obtained by the system is R. This is usually called “return
function”:
R
a0
ss ¼ E r t þ 1 jS t ¼ S; a t ¼ a; s t þ 1 ¼ s
0
f
g
ð3:3Þ
⑤ V is criterion function (or objective function). Common criterion functions include:
expected total return during a limited period, expected discounted total return and
average return. It can be a state value function or state-movement function.
Based on the characteristics of UAV cluster’s operation, UAV, battlefield’s environment, load platform, information resources and other supporting resources are
segmented in the establishment of Markov state transition model (MDP). Thereinto,
MDP model’s state variables include usability of platform and load, position of UAV,
priority of multi-tasks and residual quantity of fuel. Usability refers to whether UAV
and load are usable or not, and residual usability, whether takeoff/recycling equipment
is usable or not. Variables of position state include current position, position of
takeoff/recycling equipment and position of 4D track of task execution in sky. Priority
is used to regulate UAVs’ takeoff and landing sequence as well as priority of multiple
tasks. Residual quantity of fuel can be classified into multiple levels to judge the order
of UAVs’ landing and priority of tasks. Finally, MDP model’s state space can be
obtained:
32
Y. Mao and Y. Hu
Précédent

- 44/679

Suivant