Snapshot Navigation in the Wavelet Domain
247
of wavelets, for both single and multi-snapshot navigation, to a series of pixelbased approaches including low-resolution features and skyline representations.
We show that wavelets improve navigational performance suggesting promise for
future, computationally efficient, navigation models.
2 Methods
In order to investigate frequency based navigation, we utilised seven models of
snapshot based route navigation, each of which operated on a different representation of images. The first stage of navigation is for the agent to traverse
a pre-defined route (in ants this would be guided by path integration) during
which the information needed for later visual navigation is stored. Here we collect images periodically along the route, and use these, after image processing,
which is the key part of the investigation, as the stored snapshots.
Models: In order to retrieve a movement direction on a memorized route, a
model M maps an image I Y to T Y under different rotations r, compares it to
a stored snapshot T X , and returns the rotation ˆ
r that minimises the Rotational
Image Difference Function (RIDF) ξ (similar to [36]). In this way a snapshot
can be used to recall the direction the agent was facing when the snapshot was
stored, also known as a visual compass. We measure performance by calculating the absolute angular error |ˆ r − r
∗
| between the angle ˆ
r that minimises the
RIDF and the true bearing r
∗ , which is known for each location. Small values
correspond to small errors and can be interpreted as robust orientation recovery.
The most simple model M px is directly operating in pixel space [2]. Comparison and storage are both realised in pixel space, such that the resulting RIDF
is given by
ξ px (I X , I Y , r) =
1
H ∗ W
H
h
W
w
(I X [w, h] − I
r
Y [w, h])
2
(1)
where I x [w, h] is a pixel value at width w and height h of the snapshot image,
I
r
y [w, h] is the pixel value of the current view rotated by r, H is the number of
rows and W is the number of columns.
Next, we used three wavelet based models in order to analyse whether and
how spatial localization of frequencies can be useful for orientation recovery.
Specifically, we used the discrete wavelet transform (DWT) [26]. In order to
perform a DWT on an image, it is treated as a 2D signal, where rows and
columns are processed separately. Hence, high-pass filter responses at different
orientations represent horizontal, vertical and diagonal details (edges), while
the low-pass response remains an image approximation (blurred version of the
original image). This approximation can be used as input for the next level
of filter banks. Each of our three wavelet based models M
L
∗
wv extracts detail
coefficients for each orientation (vertical, horizontal, diagonal) for a single level
of interest L
∗ by performing a Haar [10] wavelet based DWT on the image I
up to level L
∗ . It then omits all coefficients for which L < L
∗ and sets the
247
of wavelets, for both single and multi-snapshot navigation, to a series of pixelbased approaches including low-resolution features and skyline representations.
We show that wavelets improve navigational performance suggesting promise for
future, computationally efficient, navigation models.
2 Methods
In order to investigate frequency based navigation, we utilised seven models of
snapshot based route navigation, each of which operated on a different representation of images. The first stage of navigation is for the agent to traverse
a pre-defined route (in ants this would be guided by path integration) during
which the information needed for later visual navigation is stored. Here we collect images periodically along the route, and use these, after image processing,
which is the key part of the investigation, as the stored snapshots.
Models: In order to retrieve a movement direction on a memorized route, a
model M maps an image I Y to T Y under different rotations r, compares it to
a stored snapshot T X , and returns the rotation ˆ
r that minimises the Rotational
Image Difference Function (RIDF) ξ (similar to [36]). In this way a snapshot
can be used to recall the direction the agent was facing when the snapshot was
stored, also known as a visual compass. We measure performance by calculating the absolute angular error |ˆ r − r
∗
| between the angle ˆ
r that minimises the
RIDF and the true bearing r
∗ , which is known for each location. Small values
correspond to small errors and can be interpreted as robust orientation recovery.
The most simple model M px is directly operating in pixel space [2]. Comparison and storage are both realised in pixel space, such that the resulting RIDF
is given by
ξ px (I X , I Y , r) =
1
H ∗ W
H
h
W
w
(I X [w, h] − I
r
Y [w, h])
2
(1)
where I x [w, h] is a pixel value at width w and height h of the snapshot image,
I
r
y [w, h] is the pixel value of the current view rotated by r, H is the number of
rows and W is the number of columns.
Next, we used three wavelet based models in order to analyse whether and
how spatial localization of frequencies can be useful for orientation recovery.
Specifically, we used the discrete wavelet transform (DWT) [26]. In order to
perform a DWT on an image, it is treated as a 2D signal, where rows and
columns are processed separately. Hence, high-pass filter responses at different
orientations represent horizontal, vertical and diagonal details (edges), while
the low-pass response remains an image approximation (blurred version of the
original image). This approximation can be used as input for the next level
of filter banks. Each of our three wavelet based models M
L
∗
wv extracts detail
coefficients for each orientation (vertical, horizontal, diagonal) for a single level
of interest L
∗ by performing a Haar [10] wavelet based DWT on the image I
up to level L
∗ . It then omits all coefficients for which L < L
∗ and sets the
