Reinforcement Learning of Flight Behaviors
231
For use in biologically realistic flight modeling, however, simulators like these,
suffer from two important limitations. First, video game engines are designed to
run in real time for human interaction, making them orders of magnitude too
slow to collect sufficient samples for most reinforcement learning algorithms.
Second, the use of photo-realistic data rendered at a standard video game frame
rate (60–120 fps) makes them suitable for modeling data acquisition by actual
video cameras, but unsuitable for the low-resolution, fast/asynchronous visual
sensing typical of insects and other organisms [8].
2 OpenAI Gym Environment
OpenAI Gym [2] is a toolkit for developing and comparing reinforcement learning algorithms. In addition to providing a variety of reinforcement learning environments (tasks) like Atari games, it defines a simple Application Programming
Interface (API) for developing new environments. Our simulator implements this
API as follows:
– reset() Initializes the vehicle’s state vector (position and velocity) to zero.
– step() Updates the state using the dynamics equations in [1].
– render() Shows the vehicle state using a Heads-Up-Display (HUD) or thirdperson view.
In the remainder of this extended abstract, we discuss related projects based
on OpenAI Gym and provide a brief overview of work-in-progress on a new
simulator designed to address these issues in that framework.
3 Related Work
Ours is one of a very small number of published, open-source projects using
OpenAI Gym to learn flight-control mechanisms. Two others of note are (1)
Gym-FC, which uses OpenAI Gym and DRL to tune attitude controllers for
actual vehicles [6], and (2) Gym-Quad, which has been successfully used to learn
landing behaviors with spiking neural nets using a single degree of freedom of
control [4]. Because we wished to explore behaviors beyond attitude-control and
landing, we found it more straightforward to build our own Gym environment,
rather than attempting to modify these already substantial projects.
4 Results to Date
Speedup
Without calling OpenAI Gym’s optional render() function, we are able to
achieve update rates of around 28 kHz on an ordinary desktop computer –
an order of magnitude faster than the 1 kHz rate reported for AirSim [10],
and approximately three times fast as the rate we observed with our own
UnrealEngine-based simulator.
Précédent

- 246/443

Suivant