How to Reduce Computation Time
73
General Information. Similarly to the Rmax algorithm [13], we initialized the
action-values to a value of 1 so to help exploration of non-previously selected
actions, since the action-values are updated according to the previous ones.
For the MF expert, we conducted a grid search to find the best parameter-set,
i.e. parameters maximizing the total accumulated reward over a fixed duration
of 1600 timesteps. We experimentally found that the duration of 1600 actions
is a good trade-off between the time needed by the MF to start learning and a
reasonable experiment time (1600 actions correspond to about 5 hours of real
experiment). We found α = 0.6, γ = 0.9 and τ = 0.02. For the MB expert, we
chose γ = 0.95 and N = 6. For the MB expert and the MC, we chose the same
value of τ as the MF expert.
2.3 The Experimental Task
We evaluated our cognitive architecture in a navigation task. Since running 1600
actions on the robot takes about six hours, we have created a simulation of the
task where the probabilities of transitions are derived from a 13 h exploration of
the real arena (without the reward). This simulation allowed us to quickly test
multiple coordination criteria and parameterizations, before evaluating them on
a real robot.
We used 2.6 m 9.5 m arena containing obstacles (Fig. 2), and a turtlebot.
The computer uses ROS [16] to process the signals from its sensors, controls the
mobile base and interfaces with our architecture. A Kinect-1 sensor returns an
estimate of distance to obstacles in its field of view, completed by contact sensors at the front and sides of the mobile base. The robot localizes itself using the
gmapping Simultaneous Location and Mapping Algorithm (SLAM, [17]). During
a preliminary environmental exploration phase, the robot incrementally builds
a topological map by adding evenly spaced centers, and thus autonomously creating new Markovian states (Fig. 2. A). The current state (of the corresponding
MDP) is the closest center from the robot when its previous action is completed
and it evaluates the consequences. We chose to build this map beforehand and
to reuse it for each of the learning experiments, so as to reduce the sources of
behavioral variability. However, note that with the present method the system
could start with an empty map and build it incrementally, and that a new map
could be used for each experiment.
In this experiment, the robot must learn to reach a specific state of the
environment (state 18 – see Fig. 2. B). When it succeeds, it receives a unitary
reward and is randomly returned to one of the two initial positions, located in
the extremities of the arena (states 0 and 32), to start over. The goal of the
robot is first to reach state 18. The experiment involves a stable period where
the environment and reward do not change (i.e., until action 1600), followed by
a task change where the reward is moved from state 18 to state 34. We also
made a second series of experiments where the reward is fixed but obstacles
are introduced in the environment. As the state space, the action space is a
discrete space. Here, performing an action consists of going forward along 8
equally distributed allocentric directions (Fig. 2. A). As long as the robot has not
73
General Information. Similarly to the Rmax algorithm [13], we initialized the
action-values to a value of 1 so to help exploration of non-previously selected
actions, since the action-values are updated according to the previous ones.
For the MF expert, we conducted a grid search to find the best parameter-set,
i.e. parameters maximizing the total accumulated reward over a fixed duration
of 1600 timesteps. We experimentally found that the duration of 1600 actions
is a good trade-off between the time needed by the MF to start learning and a
reasonable experiment time (1600 actions correspond to about 5 hours of real
experiment). We found α = 0.6, γ = 0.9 and τ = 0.02. For the MB expert, we
chose γ = 0.95 and N = 6. For the MB expert and the MC, we chose the same
value of τ as the MF expert.
2.3 The Experimental Task
We evaluated our cognitive architecture in a navigation task. Since running 1600
actions on the robot takes about six hours, we have created a simulation of the
task where the probabilities of transitions are derived from a 13 h exploration of
the real arena (without the reward). This simulation allowed us to quickly test
multiple coordination criteria and parameterizations, before evaluating them on
a real robot.
We used 2.6 m 9.5 m arena containing obstacles (Fig. 2), and a turtlebot.
The computer uses ROS [16] to process the signals from its sensors, controls the
mobile base and interfaces with our architecture. A Kinect-1 sensor returns an
estimate of distance to obstacles in its field of view, completed by contact sensors at the front and sides of the mobile base. The robot localizes itself using the
gmapping Simultaneous Location and Mapping Algorithm (SLAM, [17]). During
a preliminary environmental exploration phase, the robot incrementally builds
a topological map by adding evenly spaced centers, and thus autonomously creating new Markovian states (Fig. 2. A). The current state (of the corresponding
MDP) is the closest center from the robot when its previous action is completed
and it evaluates the consequences. We chose to build this map beforehand and
to reuse it for each of the learning experiments, so as to reduce the sources of
behavioral variability. However, note that with the present method the system
could start with an empty map and build it incrementally, and that a new map
could be used for each experiment.
In this experiment, the robot must learn to reach a specific state of the
environment (state 18 – see Fig. 2. B). When it succeeds, it receives a unitary
reward and is randomly returned to one of the two initial positions, located in
the extremities of the arena (states 0 and 32), to start over. The goal of the
robot is first to reach state 18. The experiment involves a stable period where
the environment and reward do not change (i.e., until action 1600), followed by
a task change where the reward is moved from state 18 to state 34. We also
made a second series of experiments where the reward is fixed but obstacles
are introduced in the environment. As the state space, the action space is a
discrete space. Here, performing an action consists of going forward along 8
equally distributed allocentric directions (Fig. 2. A). As long as the robot has not
