38
X-Machines for Agent-Based Modeling: FLAME Perspectives
be categorized into specific disciplines - deliberative or reactive [205]. Sloman’s [182] work supported use of hybrid architectures, with SimAgent using a detailed agent architecture to encompass most attributes (Figure 2.4).
Wooldridge [204] studied developing computational logics behind architectures
of multi-agent systems.
2.5 Mathematical Foundations
Multi-agent systems are temporal systems that are highly dependent on
time steps. This allows agents to have a set of allowable action set, A t , made
available depending on current time step. Day [49] discussed that these actions
chosen by agents at time t + 1, a t+1 , are dependent on various stimuli. These
are defined as follows:
a t+1 = f (o t , m t , d t , x t , u t )
(2.1)
where
• Observation of agent at time t + 1, o t+1 = σ(a t , s t ), where s t is environment state at time t.
• Memory of agent at time t + 1, m t+1 = µ(o t+1 , a t , s t ).
• A process of agent at time t + 1, d t+1 = π(m t+1 , a t , s t ).
• Plan of agent at time t + 1, x t+1 = δ(d t+1 , a t , s t ).
• Implementation of agent at time t + 1, u t+1 = ι(x t+1 , a t , s t ).
The action structure is very explicitly produced by modelers or programmers, as a step-by-step procedure when creating predictable agents. An aspect
ignored above is the learning capability of the agent. Most actions may not
be chosen during a simulation. The agent should be able to evaluate available
actions and modify them to suit its purpose. This process of learning encourages the agent to optimize its behavior, to better suit the conditions at time
t. To enable this, the agent code needs a feedback to assess its performance.
Machine learning techniques have used various methods to construct optimizing of artificial agents. Reinforcement learning can allow agents to optimize
themselves in dynamic environments. To achieve this, agents have a method
to assess their performance in certain situations using a reward structure.
1. At time t, agent sees the environment state, s t ∈ S and set of possible
actions at this state, A(s t ). Note, previously a set of allowable actions
were dependent on time A t . Now, this is dependent on the environment’s
state, bringing in awareness of the agent’s surroundings, A(s t ).
Précédent

- 67/329

Suivant