72
R. Dromnelle et al.
R(s) is the instant reward received for reaching the state s and γ the decay
rate of future rewards and the s
the state reached after executing a.
Inference Process. Since the MF expert does not use planning, its inference
process consists only in reading from the table that contains all the action-values
the one that corresponds to performing the action a in the state s.
Decision Process. The decision process is the same as for the MB expert (3).
Meta-controller and Arbitration Method. The MC is in charge of selecting
which expert will generate the behavior using a novel arbitration criterion which
is a trade-off between the quality of learning and the cost of inference.
Quality of Learning. At each time step t, if the expert E is selected to lead the
decision, its action selection probabilities (3) are filtered using a low-pass filter
and stored by the system :
f (P (a|s, E, t)) = (1.0 − α) ∗ f (P (a|s, E, t − 1)) + α ∗ P (a|s, E, t)
(5)
Else, no low-pass filter is applied:
f (P (a|s, E, t)) = f (P (a|s, E, t − 1))
Using the filtered action probability distribution f (P (a|s, E, t)), the MC can
compute the entropy H(s, E, t) of each expert, which has previously been found
to reflect the quality of learning in humans [14]:
H(s, E, t) = −
|A|
a=0
f (P (a|s, E, t)) · log 2 (f (P (a|s, E, t)))
(6)
Cost of Inference. At each time step t and for each state s, the duration T (s, E, t)
of the inference process of each expert is recorded and filtered in the same way
as the action selection probabilities (5).
Arbitration Criterion. Using the quality of learning and the cost of inference,
the MC computes one expert-value Q(s, E, t) for each expert:
Q(s, E, t) = −(H(s, E, t) + κT (s, E, t))
(7)
where κ = e
−ηH(s,M F,t) allows to weight the impact of time in the criterion:
The lower the entropy of the distribution of MF action probabilities, the more
weight the time taken to perform the inference process has in the equation. η is
a constant parameter (η = 7) weighting the entropy term and set according to
a Pareto front analysis [15] (not shown here). We were looking for a κ that minimizes the cost of inference, while maximizing the agent’s ability to accumulate
reward over time.
Finally, the MC converts the estimation of expert-values Q(s, E, t) into a
distribution of expert probabilities using a softmax function (3), and draws the
expert proposal from this distribution. The inference process of the unchosen
expert is inhibited, which thus allows the system to save computation time.
R. Dromnelle et al.
R(s) is the instant reward received for reaching the state s and γ the decay
rate of future rewards and the s
the state reached after executing a.
Inference Process. Since the MF expert does not use planning, its inference
process consists only in reading from the table that contains all the action-values
the one that corresponds to performing the action a in the state s.
Decision Process. The decision process is the same as for the MB expert (3).
Meta-controller and Arbitration Method. The MC is in charge of selecting
which expert will generate the behavior using a novel arbitration criterion which
is a trade-off between the quality of learning and the cost of inference.
Quality of Learning. At each time step t, if the expert E is selected to lead the
decision, its action selection probabilities (3) are filtered using a low-pass filter
and stored by the system :
f (P (a|s, E, t)) = (1.0 − α) ∗ f (P (a|s, E, t − 1)) + α ∗ P (a|s, E, t)
(5)
Else, no low-pass filter is applied:
f (P (a|s, E, t)) = f (P (a|s, E, t − 1))
Using the filtered action probability distribution f (P (a|s, E, t)), the MC can
compute the entropy H(s, E, t) of each expert, which has previously been found
to reflect the quality of learning in humans [14]:
H(s, E, t) = −
|A|
a=0
f (P (a|s, E, t)) · log 2 (f (P (a|s, E, t)))
(6)
Cost of Inference. At each time step t and for each state s, the duration T (s, E, t)
of the inference process of each expert is recorded and filtered in the same way
as the action selection probabilities (5).
Arbitration Criterion. Using the quality of learning and the cost of inference,
the MC computes one expert-value Q(s, E, t) for each expert:
Q(s, E, t) = −(H(s, E, t) + κT (s, E, t))
(7)
where κ = e
−ηH(s,M F,t) allows to weight the impact of time in the criterion:
The lower the entropy of the distribution of MF action probabilities, the more
weight the time taken to perform the inference process has in the equation. η is
a constant parameter (η = 7) weighting the entropy term and set according to
a Pareto front analysis [15] (not shown here). We were looking for a κ that minimizes the cost of inference, while maximizing the agent’s ability to accumulate
reward over time.
Finally, the MC converts the estimation of expert-values Q(s, E, t) into a
distribution of expert probabilities using a softmax function (3), and draws the
expert proposal from this distribution. The inference process of the unchosen
expert is inhibited, which thus allows the system to save computation time.
