180
S. Iacob et al.
[9], and the learning mechanism used is Spike Timing-Dependent Plasticity. The
SNN model is built with Izhikevich neurons [18], and consists of 7 input populations and 4 output populations. The input populations encode proprioceptive
feedback (i.e. joint position) and desired cartesian coordinates (i.e. TCP position), whereas the output populations encode the desired motor commands. The
input and output populations are connected in an all-to-all manner with STDP
synapses. During the motor-babbling phase (i.e. the training phase), the output
population activates according to randomly generated motor commands, whereas
the input layer activates according to the corresponding desired TCP position,
which is computed using forward kinematics equations. By updating the synapse
weights according to the STDP learning rule, this model can autonomously learn
the inverse kinematics transformations. On the other hand, using a predefined
forward kinematics equation to compute goal states corresponding to the generated movements means that this model is also not able to adapt to changes in
arm parameters after the training phase. Although the motor babbling learning
phase can be viewed as on-line learning, the learning process is stopped in the
following ‘performance phase’, where desired TCP position is fed through the
network and the motor commands are given as output.
REACH sits somewhere in between the above mentioned models in terms
of plasticity. It is a SNN designed for target reaching with a 2D arm simulation. Most motor equations are pre-defined and estimated during build-time of
the network by Nengo, similar to the model by Tieck et al. However, REACH
includes an adaptive component: the cerebellum (CB). Similar to its biological
counterpart, the REACH CB compensates perturbations and errors in the arm
movements. This error compensation is learned on-line with the PES learning
rule, using an efferent copy of the torque signal as error signal. REACH was
shown to perform well in the 8-reach task, shows human-like reaching movements, and is able learn long-term adaptations to force field perturbations faster
than humans. In our view, REACH makes good use of pre-programmed domain
knowledge of kinematics, as well as an adaptive component that learns on-line.
3 Methods
Our modified implementation of REACH is based on the original movement
equation [10].
u = ˆ
J
T M x u x − K v M ˙
q + u adapt
(1)
Here, u is the resulting joint torque output, u x is the desired force in hand
space, ˆ
J is the Jacobian, indicating how change in joint angles affects hand
position, M x the inertia matrix in hand space, the term K v M ˙
q accounts for the
arm’s current velocity (M: inertia matrix in joint space, K v : constant, and ˙
q:
joint velocity), and u adapt represents the adaptive compensation learned by the
cerebellum (CB) network.
In essence, Eq. 1 remains the same, since it does not depend on the number of joint dimenensions and cartesian dimensions. However, the individual
components are changed to account for the higher dimensionality and different
S. Iacob et al.
[9], and the learning mechanism used is Spike Timing-Dependent Plasticity. The
SNN model is built with Izhikevich neurons [18], and consists of 7 input populations and 4 output populations. The input populations encode proprioceptive
feedback (i.e. joint position) and desired cartesian coordinates (i.e. TCP position), whereas the output populations encode the desired motor commands. The
input and output populations are connected in an all-to-all manner with STDP
synapses. During the motor-babbling phase (i.e. the training phase), the output
population activates according to randomly generated motor commands, whereas
the input layer activates according to the corresponding desired TCP position,
which is computed using forward kinematics equations. By updating the synapse
weights according to the STDP learning rule, this model can autonomously learn
the inverse kinematics transformations. On the other hand, using a predefined
forward kinematics equation to compute goal states corresponding to the generated movements means that this model is also not able to adapt to changes in
arm parameters after the training phase. Although the motor babbling learning
phase can be viewed as on-line learning, the learning process is stopped in the
following ‘performance phase’, where desired TCP position is fed through the
network and the motor commands are given as output.
REACH sits somewhere in between the above mentioned models in terms
of plasticity. It is a SNN designed for target reaching with a 2D arm simulation. Most motor equations are pre-defined and estimated during build-time of
the network by Nengo, similar to the model by Tieck et al. However, REACH
includes an adaptive component: the cerebellum (CB). Similar to its biological
counterpart, the REACH CB compensates perturbations and errors in the arm
movements. This error compensation is learned on-line with the PES learning
rule, using an efferent copy of the torque signal as error signal. REACH was
shown to perform well in the 8-reach task, shows human-like reaching movements, and is able learn long-term adaptations to force field perturbations faster
than humans. In our view, REACH makes good use of pre-programmed domain
knowledge of kinematics, as well as an adaptive component that learns on-line.
3 Methods
Our modified implementation of REACH is based on the original movement
equation [10].
u = ˆ
J
T M x u x − K v M ˙
q + u adapt
(1)
Here, u is the resulting joint torque output, u x is the desired force in hand
space, ˆ
J is the Jacobian, indicating how change in joint angles affects hand
position, M x the inertia matrix in hand space, the term K v M ˙
q accounts for the
arm’s current velocity (M: inertia matrix in joint space, K v : constant, and ˙
q:
joint velocity), and u adapt represents the adaptive compensation learned by the
cerebellum (CB) network.
In essence, Eq. 1 remains the same, since it does not depend on the number of joint dimenensions and cartesian dimensions. However, the individual
components are changed to account for the higher dimensionality and different
