76
R. Dromnelle et al.
and only updates its action-values locally: a method that takes longer to be
effective. Finally, we can observe that the DQN agent learns and adapts less well
than all other agents. As it is a model-free algorithm, it is not surprising that
agents using the MB expert are more efficient and adaptive. The DQN is also
worse than our tabular MF because it has much more memorized values (i.e. the
weights of the network) to adapt before being able to provide correct outputs:
the training of deep neural networks generally require several hundred thousand
iterations. Such number are much too large, when targeting applications to real
robot experiments, where learning on-the-fly is required. Replay mechanisms,
or training in simulation, could be used to speed-up learning of the DQN, but
these additional computations would clearly increase the computational cost of
the resulting system.
Unsurprisingly, the MF only agent has a very low computational cumulated
cost (Fig. 3. B) since its inference process simply consists in reading from the
table that contains all the actions-values, while the MB only agent has a high
computational cost, because its inference process is a planning method. The MCEC agent, which exhibits a performance similar to the MB, has a computational
cost three times smaller: the average cumulative time at the end of the experiment spent by the MB only agent on its inference process is1750 s versus 500s
for the MC-EC agent at action 6400. It is to be noted that the meta-controller
has in any case a very low cost, similar to the MF expert, of 10
− 5 seconds per
iteration on average. In this system, only the MB expert is expensive, with an
average cost of 10
− 2 seconds per iteration. The cost of using a meta-controller
is therefore negligible compared to what it brings in terms of overall savings.
The dynamics of the selection of the experts by the MC, expressed in terms
of selection probabilities (Fig. 3. C), displays three different phases:
The MF Exploring Phase (1 on Fig. 3. C). Before the discovery of the
position of the reward, the agent uses mainly the MF expert. This is due to the
difference in the method for updating action-values between the two experts.
With the same initial values and the set of parameters we have defined, the
action-values of the MF expert decrease slightly more than those of the MB
expert, which drives a more pronounced decrease of the entropy of the action
probability distribution. In addition, since we do not have an expert specialized
in exploration, it makes sense to use the cheapest expert until the position of
the reward has been discovered. About exploration, other studies propose to deal
between three experts: a MB expert, a MF expert and an expert specialized in
the exploration of the environment [3].
The MB Driving Phase (2 on Fig. 3. C). After finding the first reward
the MB expert progressively takes the lead on the decision because its process
of inference needs only to find the reward once to spread action-values into its
transition model. It finds the reward more easily than the MF expert, and so,
its performance increases.
The MF Driving Phase (3 on Fig. 3. C). The MF expert learns by demonstration from the MB expert, and thus spreads action-values from state to state
R. Dromnelle et al.
and only updates its action-values locally: a method that takes longer to be
effective. Finally, we can observe that the DQN agent learns and adapts less well
than all other agents. As it is a model-free algorithm, it is not surprising that
agents using the MB expert are more efficient and adaptive. The DQN is also
worse than our tabular MF because it has much more memorized values (i.e. the
weights of the network) to adapt before being able to provide correct outputs:
the training of deep neural networks generally require several hundred thousand
iterations. Such number are much too large, when targeting applications to real
robot experiments, where learning on-the-fly is required. Replay mechanisms,
or training in simulation, could be used to speed-up learning of the DQN, but
these additional computations would clearly increase the computational cost of
the resulting system.
Unsurprisingly, the MF only agent has a very low computational cumulated
cost (Fig. 3. B) since its inference process simply consists in reading from the
table that contains all the actions-values, while the MB only agent has a high
computational cost, because its inference process is a planning method. The MCEC agent, which exhibits a performance similar to the MB, has a computational
cost three times smaller: the average cumulative time at the end of the experiment spent by the MB only agent on its inference process is1750 s versus 500s
for the MC-EC agent at action 6400. It is to be noted that the meta-controller
has in any case a very low cost, similar to the MF expert, of 10
− 5 seconds per
iteration on average. In this system, only the MB expert is expensive, with an
average cost of 10
− 2 seconds per iteration. The cost of using a meta-controller
is therefore negligible compared to what it brings in terms of overall savings.
The dynamics of the selection of the experts by the MC, expressed in terms
of selection probabilities (Fig. 3. C), displays three different phases:
The MF Exploring Phase (1 on Fig. 3. C). Before the discovery of the
position of the reward, the agent uses mainly the MF expert. This is due to the
difference in the method for updating action-values between the two experts.
With the same initial values and the set of parameters we have defined, the
action-values of the MF expert decrease slightly more than those of the MB
expert, which drives a more pronounced decrease of the entropy of the action
probability distribution. In addition, since we do not have an expert specialized
in exploration, it makes sense to use the cheapest expert until the position of
the reward has been discovered. About exploration, other studies propose to deal
between three experts: a MB expert, a MF expert and an expert specialized in
the exploration of the environment [3].
The MB Driving Phase (2 on Fig. 3. C). After finding the first reward
the MB expert progressively takes the lead on the decision because its process
of inference needs only to find the reward once to spread action-values into its
transition model. It finds the reward more easily than the MF expert, and so,
its performance increases.
The MF Driving Phase (3 on Fig. 3. C). The MF expert learns by demonstration from the MB expert, and thus spreads action-values from state to state
