23 Entity-Event Ontology Construction by Conceptualization …
349
Fig. 23.1 High-level architecture of computational linguistics subsystem
matching, and conditional random fields. The latter method is superior to other
ones like maximum-entropy Markov models because it has no label bias and does
not require modeling dependencies among observables. In addition to NLP tools,
analytical modules are used: culture detection, ontological part of entity resolution,
and so on.
Eventually, all events are labeled with type, temporal information, entities
involved, and which event attribute they belong to, related entities that are not
involved, mentions and their sentiment, source, and external context.
Our system is organized in layers as it is shown in Fig. 23.2. Each layer consists of
a set of microservices communicating through a messaging system. Web harvester
subsystem is responsible for mining text data from the Internet. It takes a list of
sources and robots as input and produces unstructured texts free of unwanted content
and context: time of access, time of publication according to the source, author, URL,
etc.
Computational linguistics subsystem takes as purified texts from the harvester
subsystem as input. Its output is a set of XML-formatted text fragments with highlighted entities, events, temporal and spatial labels, and so forth. It should be emphasized that this subsystem is the only place in the system that depends on a language.
Support of a new language is implemented by adding a new module for it in this
Précédent

- 346/374

Suivant