156 Environments of intelligence
expression and intonation to match, provided that the input is meaningful and
semantically located within a specified range of pre-defined topics that are made
known to the human counterpart.
This is, in some respects, the reverse to approaches in social robotics, where
non-verbal interaction is designed for symmetry, but where the robot may not be
able to understand or articulate sentences in natural language. The human-likeness
of social robots is primarily located on the level of behaviour and interactive abilities of the robot, and only partly on human-likeness in appearance. One of the
best-known examples of a humanoid robot designed for simulating sub-linguistic,
emotional levels of human communication is MIT’s KISMET, to be followed by
partly more advanced, partly differently disposed models such as LEONARDO,
MERZ and NEXI (the line of research to which I am referring to here was first
documented in Breazeal 2002; 2003; Brooks 2002a; 2002b, with newer developments to be discussed later in this section). It consists of a desk-mounted head
with some key human facial features that are animated by small electric motors:
mouth, ears, eyes and eyebrows, which are connected to a visual and auditory
system via a neural network. By those means, KISMET produces reactions to
people’s posture, gestures, facial expressions and prosody, but not to the content
of what is said, which to analyse it has no means. Nonetheless, KISMET’s reactions proved to be realistic and human-like enough to their human partners to
allow for a degree of natural communicative interaction that is said to have lured
not only its designers but also even avowed sceptics into ascribing beliefs, desires
and emotions to the robot in the course of interaction (see Brooks 2002a, 149).
In some respects, embodied conversational agents in particular may appear as
players in an advanced and, notably, non-blinded version of Turing’s imitation
game: similarity to human appearance and behaviour is the proximate aim and,
at least to the human counterpart, more channels are available through which to
receive information on the computer’s skills in impersonating a human expert in
some specified field of knowledge. Social robots offer a different take on imitation,
in allowing for a more natural and symmetric interaction in terms of visual and auditory cues while not, or only rudimentarily, offering a level of verbal conversation.
These different levels of interaction seemingly match the different levels of
embodiment of the agents: the physical embodiment of the social robot allows for
a more immediate, embodied kind of communication that may spare the verbal
level altogether, whereas the ‘virtual’, that is screen-bound, embodiment of the
conversational agent is more remote in interactional terms and has to rely on verbal cues on the agent’s side. Especially in the case of social robots, the embodied
aspects of the interaction include for example cues of spatial distance vs. proximity.
A human being will afford flinching to the robot when getting uncomfortably close
to it, just as the robot is capable of affording flinching to the human. Such elements
of interaction would be difficult to simulate for a screen-bound agent on a display
that is, qua being a display, delimited in the sense introduced for pictures by James
Jerome Gibson. This observation, however, does not speak against the possibility
of including a capacity for using some visual cues in virtually embodied agents, let
alone against adding advanced capacities for verbal behaviour to humanoid robots.
expression and intonation to match, provided that the input is meaningful and
semantically located within a specified range of pre-defined topics that are made
known to the human counterpart.
This is, in some respects, the reverse to approaches in social robotics, where
non-verbal interaction is designed for symmetry, but where the robot may not be
able to understand or articulate sentences in natural language. The human-likeness
of social robots is primarily located on the level of behaviour and interactive abilities of the robot, and only partly on human-likeness in appearance. One of the
best-known examples of a humanoid robot designed for simulating sub-linguistic,
emotional levels of human communication is MIT’s KISMET, to be followed by
partly more advanced, partly differently disposed models such as LEONARDO,
MERZ and NEXI (the line of research to which I am referring to here was first
documented in Breazeal 2002; 2003; Brooks 2002a; 2002b, with newer developments to be discussed later in this section). It consists of a desk-mounted head
with some key human facial features that are animated by small electric motors:
mouth, ears, eyes and eyebrows, which are connected to a visual and auditory
system via a neural network. By those means, KISMET produces reactions to
people’s posture, gestures, facial expressions and prosody, but not to the content
of what is said, which to analyse it has no means. Nonetheless, KISMET’s reactions proved to be realistic and human-like enough to their human partners to
allow for a degree of natural communicative interaction that is said to have lured
not only its designers but also even avowed sceptics into ascribing beliefs, desires
and emotions to the robot in the course of interaction (see Brooks 2002a, 149).
In some respects, embodied conversational agents in particular may appear as
players in an advanced and, notably, non-blinded version of Turing’s imitation
game: similarity to human appearance and behaviour is the proximate aim and,
at least to the human counterpart, more channels are available through which to
receive information on the computer’s skills in impersonating a human expert in
some specified field of knowledge. Social robots offer a different take on imitation,
in allowing for a more natural and symmetric interaction in terms of visual and auditory cues while not, or only rudimentarily, offering a level of verbal conversation.
These different levels of interaction seemingly match the different levels of
embodiment of the agents: the physical embodiment of the social robot allows for
a more immediate, embodied kind of communication that may spare the verbal
level altogether, whereas the ‘virtual’, that is screen-bound, embodiment of the
conversational agent is more remote in interactional terms and has to rely on verbal cues on the agent’s side. Especially in the case of social robots, the embodied
aspects of the interaction include for example cues of spatial distance vs. proximity.
A human being will afford flinching to the robot when getting uncomfortably close
to it, just as the robot is capable of affording flinching to the human. Such elements
of interaction would be difficult to simulate for a screen-bound agent on a display
that is, qua being a display, delimited in the sense introduced for pictures by James
Jerome Gibson. This observation, however, does not speak against the possibility
of including a capacity for using some visual cues in virtually embodied agents, let
alone against adding advanced capacities for verbal behaviour to humanoid robots.
