206
E. Kagioulis et al.
IDF (C(x, θ), S((y, φ)) =
M
i=1
N
j=1 (C ij (x, θ) − S ij (y, φ)) 2
M × N
(1)
where C(x, θ) is an M × N panoramic view from position x at an orientation
θ with pixel greyscale values C ij at row i and column j and S(y, φ) is a view
stored from position y at orientation φ with pixel values S ij .
The second method uses the correlation coefficient (CC) to derive the similarity between images [24] via:
CC(C(x, θ), S((y, φ)) =
M
i=1
N
j=1 (C ij − ¯
C)(S ij − ¯
S)
M
i=1
N
j=1 (C ij − ¯
C) 2
M
i=1
N
j=1 (S ij − ¯
S) 2
(2)
where ¯
C and ¯
S are the mean pixel values of C(x, θ) and S(y, φ) respectively.
To derive a heading from comparing C(x, θ) and S(y, φ), we ‘rotate’ the
current view by varying θ through 360
◦ (in steps of one degree) and use either
the IDF or CC to determine the orientation ˆ
θ at which C best matches S. For the
IDF a smaller value indicates a higher similarity while for CC it is the highest
value. In both cases the agent will move in the direction that is most familiar
thus pairing visual similarity to action.
Route Navigation Using a Sequential Memory Window: In our route
navigation algorithms, the agent first traverses a predefined route (e.g. blue
line in Fig. 1 (A)) during which it periodically stores views, or snapshots. In
the experiments in this paper, snapshots are stored every 1 cm and the facing
direction of each snapshot is set as the direction to the next snapshot (or the
direction of the penultimate snapshot in the case of the last snapshot in the
sequence). Unlike our previous work (e.g. [3]), the sequence of the snapshots
during the training route is retained, although the distance between them (and
as a corollary, the fact that they are regularly sampled) is not important. The
agent subsequently navigates by ‘rotating’ on the spot and comparing rotated
versions of the current view individually to a subset of the stored snapshots using
either the IDF or CC as described in the previous section. The direction which
results in the lowest minimum for the IDF/highest maximum for the CC across
the subset of snapshots is selected as the movement direction.
In contrast to the standard PM model [3] which compares the current view to
all snapshots, the sequential memory window model compares the current view
only to a subset of n consecutive images of the image memory. The best match
is used to determine both current movement direction and the start position of
the memory window for the next comparison. While other options are possible
for initialising the next window, this is reasonable under the assumption that the
agent moves fairly directly from start to end and that test points are well-spaced.
For the first window, as the test positions are near the start of the route, we use
the first snapshot (see Discussion for alternatives).
Précédent

- 221/443

Suivant