11
2.3 Extracting and quantifying digital trends
Secondly, the identified trends are extracted and quantified by a text-mining algorithm from
the consulting papers. This is done by on a standard machine with two cores @ 2.3 GHz and
8 GB ram. The algorithm analyzes the respective trends according to the occurrence in a single paper (Term Occurrence) and total number of words in respective paper. From this data
the nominal term frequency is calculated, which allows to compare the term frequency (TF)
from different papers with different length. The binary term occurrence (BTO) can be used
to determine the number of articles in which the respective trend was mentioned. The paper
frequency (PF) results from the total number of articles and the BTO.
Content relations between individual trends help to classify these and draw a superordinate picture of all relations of current digital technologies. The links are extracted from the
consulting paper by using the order of words in the text to draw a conclusion on which trends
are mentioned together and might have a direct connection. To increase the validity of these
connections, semantic relatedness from Wikipedia are used to complete the network. Particularly, the direct links to related topics or terms and assigned categories are of great interest
and investigated for more connections. Detailed information about the usage of Wikipedia
related text-mining is provided by Gabrilovich and Markovitch (2007) and Simanovsky and
Ulanov (2011). The algorithm were run on a 4 cores @ 3.4 GHz and 32 GB machine.
Trends and the links between the trends can be seen as a big network. The resulting complex
is further investigated in networking analyses, by focusing on the key parameters weighted
degree, centrality as well as layout analytics, which were discussed in Cherven (2015) and
Opsahl et al. (2010). The open source program Gephi is used to visualize the network and to
highlight the links between the trends in a networking structure. Nodes and edges characterize these structures. In this case, the nodes are represented by the individual trends. Edges
show the relation between two trends. (Chiarello et al. 2018). A forced-directed algorithm
(Force Atlas 2 by Jacomy et al. (2014)), which takes attraction and repulsion of the single
trends and their connections in consideration was used for the layout of the network.
As mentioned in the introduction, a reliable and extensive database is necessary to analyze
the implementation level of new digital technologies. In this case, articles from the onemine.
org database are used together with a text mining method (see Figure 1b), which filters the
paper for the identified trends. The resulting papers are then searched for all proper nouns
and, in a next step, analyzed for mine operation names. The mine and the mentioned trend
can then be linked.
3 RESULTS
3.1 Data collection from the database
The Forbes list for metals and mining consulting consists of 13 international acting companies. Eight of these companies offer 28 mining related online white papers in total. 26 papers
from seven companies address directly digital mining trends. The request on “onemine.org”
database focuses on all research papers published after 2010 in English via seven professional
mining organizations. This resulted in 2400 individual papers.
By analyzing the consulting papers, 209 digital trend terms were recognized. After eliminating duplicates and use of a uniform spelling, 107 individual terms were identified. Furthermore, the trends were divided into 21 synonym groups and finally clustered in the five
categories that are shown in Table 1.
3.2 Extracted digital trends
The text-mining algorithm took around 3 min. In a first run the term occurrence (TO) of
the single trends were calculated for every paper, where 820 mentions in all 26 papers were
detected. Figure 2 shows the top 20 trends with the highest TO in all consulting white papers.
In detail, the results show a significant peak for the term “automation” with 153 mentions.
2.3 Extracting and quantifying digital trends
Secondly, the identified trends are extracted and quantified by a text-mining algorithm from
the consulting papers. This is done by on a standard machine with two cores @ 2.3 GHz and
8 GB ram. The algorithm analyzes the respective trends according to the occurrence in a single paper (Term Occurrence) and total number of words in respective paper. From this data
the nominal term frequency is calculated, which allows to compare the term frequency (TF)
from different papers with different length. The binary term occurrence (BTO) can be used
to determine the number of articles in which the respective trend was mentioned. The paper
frequency (PF) results from the total number of articles and the BTO.
Content relations between individual trends help to classify these and draw a superordinate picture of all relations of current digital technologies. The links are extracted from the
consulting paper by using the order of words in the text to draw a conclusion on which trends
are mentioned together and might have a direct connection. To increase the validity of these
connections, semantic relatedness from Wikipedia are used to complete the network. Particularly, the direct links to related topics or terms and assigned categories are of great interest
and investigated for more connections. Detailed information about the usage of Wikipedia
related text-mining is provided by Gabrilovich and Markovitch (2007) and Simanovsky and
Ulanov (2011). The algorithm were run on a 4 cores @ 3.4 GHz and 32 GB machine.
Trends and the links between the trends can be seen as a big network. The resulting complex
is further investigated in networking analyses, by focusing on the key parameters weighted
degree, centrality as well as layout analytics, which were discussed in Cherven (2015) and
Opsahl et al. (2010). The open source program Gephi is used to visualize the network and to
highlight the links between the trends in a networking structure. Nodes and edges characterize these structures. In this case, the nodes are represented by the individual trends. Edges
show the relation between two trends. (Chiarello et al. 2018). A forced-directed algorithm
(Force Atlas 2 by Jacomy et al. (2014)), which takes attraction and repulsion of the single
trends and their connections in consideration was used for the layout of the network.
As mentioned in the introduction, a reliable and extensive database is necessary to analyze
the implementation level of new digital technologies. In this case, articles from the onemine.
org database are used together with a text mining method (see Figure 1b), which filters the
paper for the identified trends. The resulting papers are then searched for all proper nouns
and, in a next step, analyzed for mine operation names. The mine and the mentioned trend
can then be linked.
3 RESULTS
3.1 Data collection from the database
The Forbes list for metals and mining consulting consists of 13 international acting companies. Eight of these companies offer 28 mining related online white papers in total. 26 papers
from seven companies address directly digital mining trends. The request on “onemine.org”
database focuses on all research papers published after 2010 in English via seven professional
mining organizations. This resulted in 2400 individual papers.
By analyzing the consulting papers, 209 digital trend terms were recognized. After eliminating duplicates and use of a uniform spelling, 107 individual terms were identified. Furthermore, the trends were divided into 21 synonym groups and finally clustered in the five
categories that are shown in Table 1.
3.2 Extracted digital trends
The text-mining algorithm took around 3 min. In a first run the term occurrence (TO) of
the single trends were calculated for every paper, where 820 mentions in all 26 papers were
detected. Figure 2 shows the top 20 trends with the highest TO in all consulting white papers.
In detail, the results show a significant peak for the term “automation” with 153 mentions.
