78
R. Dromnelle et al.
We obtained similar strategy alternations with the environment change consisting of obstacles introduction without moving the reward. We also observed
that geographical patterns of coordination of experts emerged over time. These
results won’t be presented in details here because of space limitations.
4 Discussion
We analyzed the behavior of a three-layered robotic architecture integrating
neuro-inspired mechanisms for the coordination of MB and MF reinforcement
learning. The novelty relies in the explicit online measure of performance and
cost of each system, so as to give control to the system with best current trade-off
between the two. We presented real and simulated navigation results in a complex, non-stationary indoor environment. The arbitration criterion proposed in
this work allowed the robot to autonomously determine when to shift between
systems during learning, generating a coherent temporal decision-making pattern that alternates between strategies over time. This promoted more flexibility
than pure MF control in response to task changes, and permitted to reach the
same level of performance than pure MB control, while dividing computation
time by three. The comparison with DQN showed that using end-to-end RL has
a computational cost not compatible with robotic constraints, and that thus
building and using a data representation adapted to the task at hand reduces
the burden on the RL part of the system, allowing for low-cost on-the-fly learning. In future work, we plan to test whether this architecture is generalizable to
other scenarios and larger spaces states, which we have already begun to do by
applying our model to a social interaction task defined by 112 states [19].
References
1. Meyer, J.-A., Guillot, A.: Biologically-inspired robots. In: Handbook of Robotics
(B. Siciliano and O. Khatib, eds.), pp. 1395–1422. Springer, Berlin (2008). https://
doi.org/10.1007/978-3-540-30301-5 61
2. Doll´ e, L., Khamassi, M., Girard, B., Guillot, A., Chavarriaga, R.: Analyzing interactions between navigation strategies using a computational model of action selection. In: International Conference on Spatial Cognition, pp. 71–86 (2008)
3. Caluwaerts, K., et al.: A biologically inspired meta-control navigation system for
the Psikharpax rat robot. Bioinspiration Biomimetics 7, 025009 (2012)
4. Zambelli, M., Demiris, Y.: Online multimodal ensemble learning using self-learned
sensorimotor representations. IEEE Trans. Cogn. Dev. Syst. 9(2), 113–126 (2016)
5. Banquet, J.-P., Hanoune, S., Gaussier, P., Quoy, M.: From cognitive to habit
behavior during navigation, through cortical-basal ganglia loops. In: Villa, A.E.P.,
Masulli, P., Pons Rivero, A.J. (eds.) ICANN 2016. LNCS, vol. 9886, pp. 238–247.
Springer, Cham (2016). https://doi.org/10.1007/978-3-319-44778-0 28
6. Lowrey, K., Rajeswaran, A., Kakade, S., Todorov, E., Mordatch, I.: Plan online,
learn offline: efficient learning and exploration via model-based control. In: International Conference on Learning Representations (2019)
Précédent

- 93/443

Suivant