74
R. Dromnelle et al.
Fig. 2. A. Map of the arena’s states. The eight-pointed star indicates the direction (in
the map) of each robot actions. B. Photo of the arena and a turtlebot heading into
the middle corridor. The state 18 (initial reward location) is represented in red. (Color
figure online)
changed state, the action is not considered as completed. However, if while the
robot moves forward, its contact sensors are activated (it bumps into a wall), then
it will move back 0.15 meters and the action is considered as completed. Finally,
according to the exact position in which the agent is located within a state, the
arrival state will not necessarily be identical for the same action performed. The
environment is therefore probabilistic, which multiplies the possibilities for the
agent. For the MB expert, this specificity implies that the transitions T (s, a, s
)
and the rewards R(s, a) are stored respectively in the model of transition T and
the model of reward R as probability distributions.
3 Results
We first present the results obtained when a virtual agent performs the task in
a simulated environment, and then, the replication of these results in the real
environment with a Turtlebot.
3.1 Simulated Task
To evaluate the performance of the virtual agent, we studied four combinations
of experts: (1) a MF only agent using only the MF expert to decide, (2) an MB
only agent using only the MB expert to decide, (3) a random coordination agent
(MC-Rnd) which coordinates the two experts randomly and (4) an Entropy and
Cost agent (MC-EC) which coordinates the two experts using the model of arbitration presented in 2.2. We also compare our agent to an agent using a reference
learning algorithm in the literature, a DQN [18]. We evaluated iteratively several
networks with various number of layers and size of layers, and selected the set of
parameters that achieved the best performance. The neural network composed
of two hidden layers of 76 neurons which takes as input a vector of size 38 (corresponding to the activity of the states, with 1 if the state is active, and 0 if
Précédent

- 89/443

Suivant