348
M. C. Ridley
In some sense, the Internet along with social networks can be considered a giant
crowdsourcing platform for a wide variety of topics. It is well-known that generalpurpose event information such as published in mainstream media can be extracted
from the Internet as well. However, it is also true for a wide range of specialized
topics like cybersecurity. For instance, 75.8% of CVE vulnerabilities related to the
Linux kernel were exposed before their official disclosure as 0-day vulnerabilities
with the average time advance of 19 days [9]. Also, 100% of NIST CVEs were also
published and described on Twitter [9]. Another example is disaster and emergency
monitoring used by agencies such as US FEMA and the UN Office for Coordination
of Humanitarian Affairs for day-to-day operations [10].
It must be noted that along with the aforementioned abundance of relevant data,
the Internet has some special traits:
• It has to be scanned regularly and in a distributed fault-tolerant manner as new
information is generated extremely fast and is very diverse.
• Storing all collected data is economically unfeasible and technically impractical,
thus requiring special tactics for getting rid of unneeded data.
• Relevant niche sources like specialized groups in social networks are important
for achieving minimal delays.
• As multiple sources and people discuss same events, their descriptions are diverse,
written in different languages, duplicated, and merged in other discussions. Also,
people might have inconsistent views on the event. Even more important, most
significant events evolve and change over time.
• All sources have different markup, styles, page organization, and so on. Robots
have to be adapted for all major sources and be smart enough to extract text with
decent quality for secondary sources. Also, some sources like social networks
provide API directly or via data provider services like GNIP.
• Extracted text usually contains banners, ads, and other disturbing content that has
to be removed.
• Each extracted text appears in some context: media type, source URL, publication date and time, and for social media—authors, likes, reposts, etc. This
context is very helpful for both end-users like professional analysts and for the
system itself. It allows analyzing bias and sentiment of sources and authors toward
different topics, country and language biases, topics with very little interest in
social networks but forced by mainstream media, negative sentiment from people
along with a positive sentiment from mainstream media, silence on some topics
by official media, and other notable cases.
As Fig. 23.1 illustrates, texts on different natural languages should be analyzed via
a set of NLP tools and analytical modules to extract entities and events and populate
ontology with them. Also, ontology is recurrently updated with information from
external sources like GeoNames which provides geographical hierarchy, DBpedia,
and WikiData. External ontologies capture useful information such as government
officials, companies and their structure, industries, and technologies.
The system uses various NLP tools like Tomita parser, StanfordNLP, OpenNLP,
and Rosette EX along with custom extractors based on regular expressions, pattern
M. C. Ridley
In some sense, the Internet along with social networks can be considered a giant
crowdsourcing platform for a wide variety of topics. It is well-known that generalpurpose event information such as published in mainstream media can be extracted
from the Internet as well. However, it is also true for a wide range of specialized
topics like cybersecurity. For instance, 75.8% of CVE vulnerabilities related to the
Linux kernel were exposed before their official disclosure as 0-day vulnerabilities
with the average time advance of 19 days [9]. Also, 100% of NIST CVEs were also
published and described on Twitter [9]. Another example is disaster and emergency
monitoring used by agencies such as US FEMA and the UN Office for Coordination
of Humanitarian Affairs for day-to-day operations [10].
It must be noted that along with the aforementioned abundance of relevant data,
the Internet has some special traits:
• It has to be scanned regularly and in a distributed fault-tolerant manner as new
information is generated extremely fast and is very diverse.
• Storing all collected data is economically unfeasible and technically impractical,
thus requiring special tactics for getting rid of unneeded data.
• Relevant niche sources like specialized groups in social networks are important
for achieving minimal delays.
• As multiple sources and people discuss same events, their descriptions are diverse,
written in different languages, duplicated, and merged in other discussions. Also,
people might have inconsistent views on the event. Even more important, most
significant events evolve and change over time.
• All sources have different markup, styles, page organization, and so on. Robots
have to be adapted for all major sources and be smart enough to extract text with
decent quality for secondary sources. Also, some sources like social networks
provide API directly or via data provider services like GNIP.
• Extracted text usually contains banners, ads, and other disturbing content that has
to be removed.
• Each extracted text appears in some context: media type, source URL, publication date and time, and for social media—authors, likes, reposts, etc. This
context is very helpful for both end-users like professional analysts and for the
system itself. It allows analyzing bias and sentiment of sources and authors toward
different topics, country and language biases, topics with very little interest in
social networks but forced by mainstream media, negative sentiment from people
along with a positive sentiment from mainstream media, silence on some topics
by official media, and other notable cases.
As Fig. 23.1 illustrates, texts on different natural languages should be analyzed via
a set of NLP tools and analytical modules to extract entities and events and populate
ontology with them. Also, ontology is recurrently updated with information from
external sources like GeoNames which provides geographical hierarchy, DBpedia,
and WikiData. External ontologies capture useful information such as government
officials, companies and their structure, industries, and technologies.
The system uses various NLP tools like Tomita parser, StanfordNLP, OpenNLP,
and Rosette EX along with custom extractors based on regular expressions, pattern
