Machine Morality
119
reinforcement learning that links these cognitive networks associated with consequentialism and deontology with two different computational learning systems,
namely model-free and model-based reinforcement learning [10].
A model-based computational system selects actions based on their presumed
outcomes while a model-free system chooses actions based on their reinforcement
history. Moreover, these computational models have clear functional mapping to
the philosophical views of deontology (model-free) and consequentialism (modelbased). As for the cognitive networks, the model-free system in the brain uses
automatic, fast and reflexive mechanisms that we are believed to be sharing
with other species. Model-based system, on the other hand, employs acquired,
controlled and reflective mechanisms that are only observed in human adults [1].
Apart from these two cognitive systems, literature suggests existence of
another distinct decision-making mechanism that we share with other animals,
the one in charge of automatic reflexive approach and withdrawal reactions to
appetitive and aversive stimuli, respectively: a so-called Pavlovian system [14].
But what role would Pavlovian system play in moral decision-making? The possible answer could be found in the changing nature of environmental conditions:
multiple decision-making modules can provide different benefits in certain contexts. Model-based control works well for simple decisions, but when the decision
tree search is computationally costly, model-free and Pavlovian systems can be
utilized to guide the search.
One promising attempt to explain the functional utility of the Pavlovian
system [10] builds upon puzzling experimental results obtained in two versions
of a classic moral dilemma (the so-called ’push-trapdoor divergence’). In this
study, a traditional trolley problem (where one was offered to pull a lever to
guide a trolley towards a single man in order to save lives of five workers) was
modified to substitute the action of pulling the lever for a physical push of the
victim. Surprisingly, this change lead to a significant difference in outcomes:
subjects were more likely to pull the lever rather than physically push a person
towards the track. The role of Pavlovian system here is suggested to be comprised
of behavioral suppression aimed at preventing aversive outcome. Such capacity
could be used, for instance, for the suppression of trains of thought or behavioral
sequences that result in aversive states, or in other words, for a pruning of the
decision tree [23]. In the context of the modified trolley setup, a physical push of
the victim would be considered a less desirable outcome based on aversive bias
induced by the Pavlovian system.
3 Computational Models of Moral Decision-Making
Recent advances coming from the Reinforcement Learning (RL) literature have
proposed several ways to combine model-free with model-based learning. Among
them, Episodic-RL and Meta-RL are the most relevant ones due to their success
in matching human data from several experimental benchmarks and improvements in learning speeds, bringing the sample-inefficiency problem of Deep RL
a bit closer to human learning timescales [8].
119
reinforcement learning that links these cognitive networks associated with consequentialism and deontology with two different computational learning systems,
namely model-free and model-based reinforcement learning [10].
A model-based computational system selects actions based on their presumed
outcomes while a model-free system chooses actions based on their reinforcement
history. Moreover, these computational models have clear functional mapping to
the philosophical views of deontology (model-free) and consequentialism (modelbased). As for the cognitive networks, the model-free system in the brain uses
automatic, fast and reflexive mechanisms that we are believed to be sharing
with other species. Model-based system, on the other hand, employs acquired,
controlled and reflective mechanisms that are only observed in human adults [1].
Apart from these two cognitive systems, literature suggests existence of
another distinct decision-making mechanism that we share with other animals,
the one in charge of automatic reflexive approach and withdrawal reactions to
appetitive and aversive stimuli, respectively: a so-called Pavlovian system [14].
But what role would Pavlovian system play in moral decision-making? The possible answer could be found in the changing nature of environmental conditions:
multiple decision-making modules can provide different benefits in certain contexts. Model-based control works well for simple decisions, but when the decision
tree search is computationally costly, model-free and Pavlovian systems can be
utilized to guide the search.
One promising attempt to explain the functional utility of the Pavlovian
system [10] builds upon puzzling experimental results obtained in two versions
of a classic moral dilemma (the so-called ’push-trapdoor divergence’). In this
study, a traditional trolley problem (where one was offered to pull a lever to
guide a trolley towards a single man in order to save lives of five workers) was
modified to substitute the action of pulling the lever for a physical push of the
victim. Surprisingly, this change lead to a significant difference in outcomes:
subjects were more likely to pull the lever rather than physically push a person
towards the track. The role of Pavlovian system here is suggested to be comprised
of behavioral suppression aimed at preventing aversive outcome. Such capacity
could be used, for instance, for the suppression of trains of thought or behavioral
sequences that result in aversive states, or in other words, for a pruning of the
decision tree [23]. In the context of the modified trolley setup, a physical push of
the victim would be considered a less desirable outcome based on aversive bias
induced by the Pavlovian system.
3 Computational Models of Moral Decision-Making
Recent advances coming from the Reinforcement Learning (RL) literature have
proposed several ways to combine model-free with model-based learning. Among
them, Episodic-RL and Meta-RL are the most relevant ones due to their success
in matching human data from several experimental benchmarks and improvements in learning speeds, bringing the sample-inefficiency problem of Deep RL
a bit closer to human learning timescales [8].
