246
S. Meyer et al.
views in memory. The rotation that minimises the difference, or is the most
familiar, yields the direction in which the ant should move [2,3,36].
Animals or robots implementing view-based homing face challenges in the
memory capacity needed to store views, as well as the computational power to
quickly compare them. For snapshot type models, the required memory increases
with route length and so does comparison time [1]. Attempts to mitigate this
effect have been made, by using different architectures of neural networks. One
example would be the use of an infomax model [2,15] to learn a familiarity
function for each route. Instead of storing all views this model was trained with
route views, such that it can output a familiarity value for any view, and use this
for orientation recovery. Storing route views implicitly in an adjustable network
like this allows for fixed memory size and comparison time up to a maximum of
route length.
Another way of decreasing memory requirements is changing the representation of visual scenes. Many view-based models operate on pixel space [8,32,36].
However, due to its anatomy the insect visual system is highly unlikely to store
and compare images in this way. Experimental evidence instead suggests that
it implements a system of visual filters [24,25,30], which process views to compressed representations formed by the filter outputs. Image differences can then
be calculated in the co-domain of this new mapping. Typical examples of such a
visual mapping include frequency based approaches like Fourier [6], Zernike [12]
and Wavelet [16] transforms.
Fourier encoded views have been used in robotics in the context of place
recognition in the past (e.g. [18]) and adapted for visual homing [27,31]. Extending on these findings, Zernike moments have been found to change between two
locations in a way that allows to derive a homing direction [29,33] and similar results have been reported for homing based on a subset of Haar wavelet
responses [14].
Fourier transforms provide information about which frequencies occur in a
signal, but do not convey any localisation of these frequencies (note, “localisation” here is used in accordance with the signal processing literature to denote
the occurrence of a frequency in sample space and not the position of an agent).
By contrast, the wavelet transform enables one to localise frequencies. This is
done by convolving the input signal with discretely shifted and scaled versions
of piecewise continuous functions called wavelets [16]. This property allows one
to create a series of filters in a sub-band coding scheme [34], where each filter
response yields a coefficient magnitude representing the occurrence of a certain
frequency at a given location. Retaining the location of visual information is
useful for navigation models (e.g. [23]).
To investigate frequency based image representations for route navigation, we
trial a sparse wavelet based representation of views for route following. We chose
filters corresponding to ‘Haar’ wavelets [10] which are plausible approximations
to filters that are implemented in the visual system of an insect. Furthermore,
they are well localised in space which for spatial tasks may be more important
than precise information in the frequency domain. We compare the performance
S. Meyer et al.
views in memory. The rotation that minimises the difference, or is the most
familiar, yields the direction in which the ant should move [2,3,36].
Animals or robots implementing view-based homing face challenges in the
memory capacity needed to store views, as well as the computational power to
quickly compare them. For snapshot type models, the required memory increases
with route length and so does comparison time [1]. Attempts to mitigate this
effect have been made, by using different architectures of neural networks. One
example would be the use of an infomax model [2,15] to learn a familiarity
function for each route. Instead of storing all views this model was trained with
route views, such that it can output a familiarity value for any view, and use this
for orientation recovery. Storing route views implicitly in an adjustable network
like this allows for fixed memory size and comparison time up to a maximum of
route length.
Another way of decreasing memory requirements is changing the representation of visual scenes. Many view-based models operate on pixel space [8,32,36].
However, due to its anatomy the insect visual system is highly unlikely to store
and compare images in this way. Experimental evidence instead suggests that
it implements a system of visual filters [24,25,30], which process views to compressed representations formed by the filter outputs. Image differences can then
be calculated in the co-domain of this new mapping. Typical examples of such a
visual mapping include frequency based approaches like Fourier [6], Zernike [12]
and Wavelet [16] transforms.
Fourier encoded views have been used in robotics in the context of place
recognition in the past (e.g. [18]) and adapted for visual homing [27,31]. Extending on these findings, Zernike moments have been found to change between two
locations in a way that allows to derive a homing direction [29,33] and similar results have been reported for homing based on a subset of Haar wavelet
responses [14].
Fourier transforms provide information about which frequencies occur in a
signal, but do not convey any localisation of these frequencies (note, “localisation” here is used in accordance with the signal processing literature to denote
the occurrence of a frequency in sample space and not the position of an agent).
By contrast, the wavelet transform enables one to localise frequencies. This is
done by convolving the input signal with discretely shifted and scaled versions
of piecewise continuous functions called wavelets [16]. This property allows one
to create a series of filters in a sub-band coding scheme [34], where each filter
response yields a coefficient magnitude representing the occurrence of a certain
frequency at a given location. Retaining the location of visual information is
useful for navigation models (e.g. [23]).
To investigate frequency based image representations for route navigation, we
trial a sparse wavelet based representation of views for route following. We chose
filters corresponding to ‘Haar’ wavelets [10] which are plausible approximations
to filters that are implemented in the visual system of an insect. Furthermore,
they are well localised in space which for spatial tasks may be more important
than precise information in the frequency domain. We compare the performance
