Machine Morality
123
First, the agent will need to acquire its Pavlovian component by engaging in
an evolutionary process that will generate a reactive control system adapted to
the specific contingencies of the world it will live in [3–5]. This process will be
necessary for the agent to develop a basic set of appetitive and aversive behaviors
that allow the agent to survive and fulfill its basics needs. A shortcut for this
process will be, naturally, either explicit or implicit integration of a fixed set of
behaviors (implementing a form of “architectural bias”).
Secondly, to train the model-free learning module, the agent will have to learn
the valence of its actions through trial and error during the course of its development. This process can occur as learning of weights in a neural network, or as
ascribing values to state-action couplets in reinforcement learning algorithms. It
will be the equivalent of the pruning process occurring in the human brain during the early stages of development. When it reaches ‘maturity’ -by completing
this developmental stage, it will be ready to build a model of the world [17]. It
is important to note that if we aim to develop a competent moral agent able to
deal with the complexities of the social world we want it to navigate, it will be
necessary to expose this agent to certain degrees of social interaction since the
very beginning. This could be done gradually, starting from dyadic interactions
in its early stages and scaling to group level scenarios later in its development.
Thirdly, after the second training stage is complete, the agent will be ready
to create internal models of the world. In this final stage of training, inspired
by the social nature of morality, the model-based learning module will need to
make a reliable model not only of its environment, but also of the other agents,
with whom it interacts [16,18,19]. Ideally, this model-based mechanism should
be able to generate internal models of other agents as well as of higher-order
entities such as a group or a population of agents. Recent work has proven
computational feasibility of such group-theory-of-mind [36].
Although the training stage of such agent might be complex, the experimental tools required for an attempt to shape such architecture already exist. For
each of these three different stages, there are specific benchmarks that we could
pursue, each of them coming from a different sub-field of AI and robotics. On
the evolutionary stage, the benchmarks can be drawn from artificial life and evolutionary robotics literature [15]. On the developmental stage, several types of
dyadic game-theoretical scenarios could be used as training examples of different
types of social interactions to test and train the agent [16,17]. For the last stage
of learning, benchmarks from the multi-agent reinforcement learning literature,
since they deal with massive multi-agent scenarios [30] in which certain norms
need to be learned to successfully complete the task [31]. Lastly, the natural
target benchmark for this proposed artificial moral agent would be to fit the
human behavioral results of the two classical versions of the Trolley Problem,
the so-called ‘push-trapdoor divergence’ [22].
In order to instantiate this approach in a concrete scenario we implemented
it in a functional industrial plant arm robot within the framework of HRRecycler project [6]. The autonomy and interaction capabilities of the industrial
arm robot itself are quite limited as it operates with a human worker through
123
First, the agent will need to acquire its Pavlovian component by engaging in
an evolutionary process that will generate a reactive control system adapted to
the specific contingencies of the world it will live in [3–5]. This process will be
necessary for the agent to develop a basic set of appetitive and aversive behaviors
that allow the agent to survive and fulfill its basics needs. A shortcut for this
process will be, naturally, either explicit or implicit integration of a fixed set of
behaviors (implementing a form of “architectural bias”).
Secondly, to train the model-free learning module, the agent will have to learn
the valence of its actions through trial and error during the course of its development. This process can occur as learning of weights in a neural network, or as
ascribing values to state-action couplets in reinforcement learning algorithms. It
will be the equivalent of the pruning process occurring in the human brain during the early stages of development. When it reaches ‘maturity’ -by completing
this developmental stage, it will be ready to build a model of the world [17]. It
is important to note that if we aim to develop a competent moral agent able to
deal with the complexities of the social world we want it to navigate, it will be
necessary to expose this agent to certain degrees of social interaction since the
very beginning. This could be done gradually, starting from dyadic interactions
in its early stages and scaling to group level scenarios later in its development.
Thirdly, after the second training stage is complete, the agent will be ready
to create internal models of the world. In this final stage of training, inspired
by the social nature of morality, the model-based learning module will need to
make a reliable model not only of its environment, but also of the other agents,
with whom it interacts [16,18,19]. Ideally, this model-based mechanism should
be able to generate internal models of other agents as well as of higher-order
entities such as a group or a population of agents. Recent work has proven
computational feasibility of such group-theory-of-mind [36].
Although the training stage of such agent might be complex, the experimental tools required for an attempt to shape such architecture already exist. For
each of these three different stages, there are specific benchmarks that we could
pursue, each of them coming from a different sub-field of AI and robotics. On
the evolutionary stage, the benchmarks can be drawn from artificial life and evolutionary robotics literature [15]. On the developmental stage, several types of
dyadic game-theoretical scenarios could be used as training examples of different
types of social interactions to test and train the agent [16,17]. For the last stage
of learning, benchmarks from the multi-agent reinforcement learning literature,
since they deal with massive multi-agent scenarios [30] in which certain norms
need to be learned to successfully complete the task [31]. Lastly, the natural
target benchmark for this proposed artificial moral agent would be to fit the
human behavioral results of the two classical versions of the Trolley Problem,
the so-called ‘push-trapdoor divergence’ [22].
In order to instantiate this approach in a concrete scenario we implemented
it in a functional industrial plant arm robot within the framework of HRRecycler project [6]. The autonomy and interaction capabilities of the industrial
arm robot itself are quite limited as it operates with a human worker through
