Agents in Economic Markets and Games
165
The game is essentially a two-player game where each player is trying to
maximize its own payoff without any consideration of what happens to the
other player.
TABLE 6.5: Prisoner sentences in PD game.
Prisoner B
Prisoner B
stays silent
betrays
Prisoner A stays silent Each serves 6 months
Prisoner A: 10 years,
Prisoner B: goes free
Prisoner A betrays
Prisoner A: goes free,
Prisoner B: 10 years
Each serves 5 years
In a one-shot game, because the players have no knowledge of other player
strategies, the game may not be very useful. But in an iterated prisoner
dilemma game, the game is repeatedly played amongst players. When repeatedly playing the game, the players have a chance to punish others, if they have
played a strategy which was unfavorable to them previously. This is similar
to reinforcement learning where, by punishing the player, they can learn the
beneficial strategies to play. The game can be repeated infinitely and eventually find an equilibrium, where they learn to play the good defect strategy to
prevent being punished in the future. In its classical form, the game presents
a Nash equilibrium when the players both defect.
Conducting the game on a trial-by-trial basis or a series of moves, the players must choose either to cooperate or defect on each trial. Table 6.6 shows
the numerical payoffs of the strategies played. Table 6.6 depicts a mathematical representation where T stands for temptation to defect, R for reward for
mutual cooperation, P for punishment for mutual defection and S for sucker’s
payoff. In this situation the following inequality will always hold:
T > R > P > S
(6.10)
TABLE 6.6: Payoff matrix in PD game, where R=3, S=0, T=5, P=1.
Cooperate Defect
Cooperate R,R (3,3)
S,T (0,5)
Defect
T,S (5,0)
P,P (1,1)
Playing the game repeatedly will eventually lead to an equilibrium, where
all players learn to defect or stay silent to achieve the maximum payoff. This
is maintained at a condition where the following rule is true [154, 48]:
2R > T + S
(6.11)
165
The game is essentially a two-player game where each player is trying to
maximize its own payoff without any consideration of what happens to the
other player.
TABLE 6.5: Prisoner sentences in PD game.
Prisoner B
Prisoner B
stays silent
betrays
Prisoner A stays silent Each serves 6 months
Prisoner A: 10 years,
Prisoner B: goes free
Prisoner A betrays
Prisoner A: goes free,
Prisoner B: 10 years
Each serves 5 years
In a one-shot game, because the players have no knowledge of other player
strategies, the game may not be very useful. But in an iterated prisoner
dilemma game, the game is repeatedly played amongst players. When repeatedly playing the game, the players have a chance to punish others, if they have
played a strategy which was unfavorable to them previously. This is similar
to reinforcement learning where, by punishing the player, they can learn the
beneficial strategies to play. The game can be repeated infinitely and eventually find an equilibrium, where they learn to play the good defect strategy to
prevent being punished in the future. In its classical form, the game presents
a Nash equilibrium when the players both defect.
Conducting the game on a trial-by-trial basis or a series of moves, the players must choose either to cooperate or defect on each trial. Table 6.6 shows
the numerical payoffs of the strategies played. Table 6.6 depicts a mathematical representation where T stands for temptation to defect, R for reward for
mutual cooperation, P for punishment for mutual defection and S for sucker’s
payoff. In this situation the following inequality will always hold:
T > R > P > S
(6.10)
TABLE 6.6: Payoff matrix in PD game, where R=3, S=0, T=5, P=1.
Cooperate Defect
Cooperate R,R (3,3)
S,T (0,5)
Defect
T,S (5,0)
P,P (1,1)
Playing the game repeatedly will eventually lead to an equilibrium, where
all players learn to defect or stay silent to achieve the maximum payoff. This
is maintained at a condition where the following rule is true [154, 48]:
2R > T + S
(6.11)
