344
M. C. Ridley
• Type of input data:
– Structured: DBpedia, relational databases, some XML.
– Unstructured: arbitrary text in one of natural of artificial languages.
– Semi-structured: unstructured data with structured parts such as Wikipedia
articles based on templates, or financial statements.
• Level of automation:
– Semi-automatic: requires user intervention.
– Automatic: after the model is ready, no control is needed.
• Learning targets:
– Concepts and instances.
– Relations: taxonomic and non-taxonomic like thematic roles and syntactic
relations.
– Axioms: used to model sentences that are true and to create new knowledge
from the existing one.
– Meta-knowledge: rules of how to learn ontology, what attributes can be
extracted.
• Purpose of the method:
– Creation of ontology from scratch.
– Ontology population and updating.
• Learning techniques:
– Linguistic: syntactic analysis, morpho-syntactic analysis, lexico-syntactic
pattern-parsing, terminology networks, syntactic frames, and text understanding techniques.
– Pattern/Template matching
1 : Hearst patterns, regular expressions, exception
templates, and symbolic interpretation rules.
– Logical: inductive logic programming, clustering, and rule learning based on
first-order logic or propositional learning.
– Statistical: hidden Markov models, sequence models, neural networks, conditional random fields, co-occurrence data, bag-of-words, and so on.
– Combined/Hybrid: various heuristics using statistical methods on top of
linguistic features. Applying methods depending on the context like WebKB
uses first-order logic rule learning along with Bayesian learning.
As this field of study is active, there are lots of tools like pioneer Text-to-Onto,
OntoLT, WebKB, DODDLE II, CRCTOL, C-Pankow, Sofie, and others based on a
1 It is worth noting that pattern matching is a common choice in information extraction tasks as it
provides very high precision, although by the cost of lower recall and constant maintenance of rules
for every supported language.
M. C. Ridley
• Type of input data:
– Structured: DBpedia, relational databases, some XML.
– Unstructured: arbitrary text in one of natural of artificial languages.
– Semi-structured: unstructured data with structured parts such as Wikipedia
articles based on templates, or financial statements.
• Level of automation:
– Semi-automatic: requires user intervention.
– Automatic: after the model is ready, no control is needed.
• Learning targets:
– Concepts and instances.
– Relations: taxonomic and non-taxonomic like thematic roles and syntactic
relations.
– Axioms: used to model sentences that are true and to create new knowledge
from the existing one.
– Meta-knowledge: rules of how to learn ontology, what attributes can be
extracted.
• Purpose of the method:
– Creation of ontology from scratch.
– Ontology population and updating.
• Learning techniques:
– Linguistic: syntactic analysis, morpho-syntactic analysis, lexico-syntactic
pattern-parsing, terminology networks, syntactic frames, and text understanding techniques.
– Pattern/Template matching
1 : Hearst patterns, regular expressions, exception
templates, and symbolic interpretation rules.
– Logical: inductive logic programming, clustering, and rule learning based on
first-order logic or propositional learning.
– Statistical: hidden Markov models, sequence models, neural networks, conditional random fields, co-occurrence data, bag-of-words, and so on.
– Combined/Hybrid: various heuristics using statistical methods on top of
linguistic features. Applying methods depending on the context like WebKB
uses first-order logic rule learning along with Bayesian learning.
As this field of study is active, there are lots of tools like pioneer Text-to-Onto,
OntoLT, WebKB, DODDLE II, CRCTOL, C-Pankow, Sofie, and others based on a
1 It is worth noting that pattern matching is a common choice in information extraction tasks as it
provides very high precision, although by the cost of lower recall and constant maintenance of rules
for every supported language.
