70
R. Dromnelle et al.
Fig. 1. The generic version of the architecture. Two experts having different properties
are computing the next action to do in the current state s. They each send monitoring
data to the meta-controller (MC) about their learning status and inference process (t1).
The MC designate the winning expert according to a criterion that uses these data
and authorizes it to carry out its inference and decision processes (t2). After making
a decision, the winning expert sends its proposition to the MC (t3), which sends the
action to the Executive Layer (t4). The effect of the executed action generates a new
perception, transformed into an abstract Markovian state, and eventually a non null
reward r, that are sent to the experts. Each expert learns according to the action
chosen by the MC, the new state reached and the reward.
actuators during interaction with the environment. The robot reaches a new
state and obtains or not a reward. The two experts use the new state and the
reward information to update their knowledge about the executed action. This
allows MB and MF experts to cooperate by learning from each others’ decision.
Compared to our previous architecture [10], several changes have been made:
The overall organization of the decision-making layer and the prioritization of
communication between modules have been changed. The MF expert is no longer
built as a neural network but as a tabular algorithm. The MC chooses which
expert is the most suitable at a given time and in a given state, and no longer
simply at a given time. And above all, we have defined a novel arbitration criterion that allows to reduce computational cost while maintaining performance.
2.2 The Decision Layer
Model-Based Expert. The MB expert learns a transition model T and a
reward model R of the problem, and uses them to compute the values of actions
in each state. These models allow to simulate over several steps the consequences
of following a given behavior and to look for desirable states to reach. Consequently, when the robot realizes that the task has changed, it can use this
knowledge of the world to instantly find the new relevant behavior. However,
Précédent

- 85/443

Suivant