How to Reduce Computation Time
77
and eventually, towards the 800th action, it reaches the performance of the MB
expert. Because the MF expert is less expensive, the model of arbitration gives
it the lead on the decision.
A MF exploring phase starts again at the 1600th action when the rewarded
state moves from state 18 to 34. Then, the MB driving and the MF driving
phases repeat.
The large standard deviation is explained by the fact that for each experiment, the agent’s strategy and behaviour can be very different, notably due to
the large number of states and possible actions, but also to the probabilistic
nature of the environment. As a result, the time of the switches from one phase
to another varied a lot from one individual to another. Nevertheless the individual behavior of each run is consistent with the average behavior presented here
(Fig. 3. D).
3.2 Real Task
Fig. 4. A. Mean performance (brown) and cost (cyan) for 100 simulated runs (dashed
curves) and 10 real runs (solid curves) of the navigation task for the MC-EC agent.
B. Mean probabilities of selection of experts by the MC using the Entropy and Cost
criterion for 10 real runs of the task. (Color figure online)
We evaluated our model of coordination on a real robot to verify that these
results cross the reality-gap. Figure 4. A compares the performance and the cost
of the MC-EC agent and the real robot (both use the same model of arbitration).
The reality gap is visible, with a drop in performance and a cost increase for
the real robot compared to the simulation. However, the model still allows the
real robot to learn and accumulate reward over time in the same way, and the
economy of cost remains advantageous.
Figure 4. B shows the dynamics of selection of the experts by the MC, for
the experiments in real environment with the real robot. Again, the three-phases
pattern is present, with only a 300 actions mean delay at the beginning of the
third phase.
77
and eventually, towards the 800th action, it reaches the performance of the MB
expert. Because the MF expert is less expensive, the model of arbitration gives
it the lead on the decision.
A MF exploring phase starts again at the 1600th action when the rewarded
state moves from state 18 to 34. Then, the MB driving and the MF driving
phases repeat.
The large standard deviation is explained by the fact that for each experiment, the agent’s strategy and behaviour can be very different, notably due to
the large number of states and possible actions, but also to the probabilistic
nature of the environment. As a result, the time of the switches from one phase
to another varied a lot from one individual to another. Nevertheless the individual behavior of each run is consistent with the average behavior presented here
(Fig. 3. D).
3.2 Real Task
Fig. 4. A. Mean performance (brown) and cost (cyan) for 100 simulated runs (dashed
curves) and 10 real runs (solid curves) of the navigation task for the MC-EC agent.
B. Mean probabilities of selection of experts by the MC using the Entropy and Cost
criterion for 10 real runs of the task. (Color figure online)
We evaluated our model of coordination on a real robot to verify that these
results cross the reality-gap. Figure 4. A compares the performance and the cost
of the MC-EC agent and the real robot (both use the same model of arbitration).
The reality gap is visible, with a drop in performance and a cost increase for
the real robot compared to the simulation. However, the model still allows the
real robot to learn and accumulate reward over time in the same way, and the
economy of cost remains advantageous.
Figure 4. B shows the dynamics of selection of the experts by the MC, for
the experiments in real environment with the real robot. Again, the three-phases
pattern is present, with only a 300 actions mean delay at the beginning of the
third phase.
