How to Reduce Computation Time
69
[3]. Another expected advantage for a robot is to detect when it can avoid the
computation time associated to a costly planning process and rely on cheaper
systems if they enable to reach the same level of performance.
In computational neuroscience, reinforcement learning (RL) algorithms have
been proposed to account for how animals initially solve a new task through
planning within a model-based (MB) system, and progressively shift to modelfree (MF) control when learning has converged [7,8]. MF learning is proposed to
represent habit learning because it takes a long time to converge, but permits fast
and efficient decisions after learning. Moreover, its slowness in learning makes it
inflexible in response to task changes, requiring that the brain switches back to
a control level similar to MB.
We have previously proposed a way to implement these principles within
a classical three-layered robot cognitive architecture, to facilitate integration
with other sensing and control components, as well as permit future transfer
to different robotic platforms [9]. Here, and after evaluating several arbitration
mechanisms between MB and MF learning systems in a previous study [10],
we present a novel one which dynamically deals between the quality of learning
and the computation cost. We test the new algorithm during simulated and real
robot navigation in a task involving paths of different lengths to the goal, deadends, and non-stationarity. We find that the algorithm flexibly and consistently
switches to MB control after environmental changes, and to MF control when the
task is stationary. Overall, the robot achieves the same performance as optimal
MB control in the task, while dividing computation time by more than two.
In summary, we propose a MB/MF algorithm using an arbitration mechanism
that coordinates the learning systems and efficiently reduces computation cost
while maintaining performance. We evaluate the algorithm both on simulated
and real robots.
2 Materials and Methods
2.1 A Robotic Architecture with a Dual Decision-Making System
The present work implements a classical three-layer robot cognitive architecture
[11,12] composed of a decision, an executive and a functional layer. The decision
layer of the proposed architecture (Fig. 1) is composed by two competing experts
which generate action propositions, each with its own method and with its own
advantages and disadvantages. These two experts are directly inspired by the
currently conventional distinction in computational neuroscience models between
goal-directed and habitual strategies [8]. The two experts run three processes in
a row: learning, inference and decision. This layer is also provided with a metacontroller (MC) in charge of arbitrating between experts. The MC determines
which expert’s proposed action will be executed in the current state, according
to an arbitration criterion.
After that, the decision layer sends the chosen action to the executive layer,
who ensures its accomplishment by recruiting robot’s skills from the functional
layer. The latter consists of a set of reactive sensorimotor loops that control
Précédent

- 84/443

Suivant