232
S. D. Levy
Learning
As a simple test of applying our platform to a reinforcement-learning task, we
used a Proximal Policy Optimization (PPO) agent [9] to learn a challenging
behavior, using a dynamics model approximating the popular DJI Phantom
quadcopter.
In this behavior, Lander2D, the agent was given the vehicle’s horizontal and
vertical coordinates, their first derivatives, and the vehicle’s roll and its first
derivative (six degrees of freedom). The control signal consisted of roll and thrust
commands (two degrees of freedom). As with the popular Lunar Lander game,
the goal was to land the vehicle between two flags in the center of the landscape,
with a large penalty (−100 points) for flying outside the landscape, a large bonus
(+100 points) for landing between the two flags, and a small penalty for distance
from the mid-point, to provide a gradient. (See Fig. 1(a) for an illustration.)
Our preliminary results were encouraging. In the Lander2D behavior, the
PPO agent learned to land the vehicle in under 9,000 training episodes (around
17 min on a Dell Optiplex 7040 eight-core desktop computer with 3.4 GHz processors). The best score achieved by the agent was competitive with the score
obtainable through a heuristic (PID control) approach, around 205 points.
Our current work involves three directions: (1) replacing the PPO agent
with a more biologically realistic spiking neural network (SNN); (2) extending
our results to a full three-dimensional simulation, as shown in Fig. 1(b); and (3)
adding a simulated vision system.
For the third direction – adding visual observation – DRL algorithms have
traditionally used convolutional neural network (CNN) layers to enable the network to learn the critical features of the environment from image pixels. CNNs
have however come under criticism for lacking biological plausibility and requiring very large amounts of computation time to learn the relationship between
the pixels and the world state [5].
To address these issues in a biologically plausible way, we are developing a
simple simulation of an event-based vision sensing [8] to incorporate into our
platform. Instead of generating simulated events by sampling of an ordinary
camera image, our event simulator uses a rudimentary perspective model to
generate pseudo-events based on the vehicle state and the position of a simulated
target object, represented as a sphere. We plan to use this simulator to model
visually-guided predation with SNNs, extending the SNN landing-behavior work
presented in [4].
5 Code
Our code, with instructions for reproducing our results, is available at https://
github.com/simondlevy/gym-copter.
Précédent

- 247/443

Suivant