310
M. Noura et al.
• Road Accident, Vehicular Accident, Traffic Accidents, Road Safety, Car Accident
Prevention, Accident Rescue Mission.
• Intersection Assistance.
• Vehicle Context-aware Services.
• Pedestrian Detection.
• Car Pooling Recommendation System.
For the set of scientific publications, we focused on the following criteria:
• Are ontology URLs available within the scientific article? Frequently, URLs are
missing. Authors have been contacted to retrieve ontology code and we enriched
the dataset when receiving positive answers. Table 1 summarizes the 16 ontologies
that share their ontologies online, which is the smart vehicle ontology dataset later
analyzed. Other ontology-based projects are referenced in Table 2, unfortunately,
the ontologies cannot be processed yet since they are not accessible.
• Are sensors mentioned within the paper?
• Are there reasoning mechanisms and already defined rules to interpret data generated by the smart vehicle applications?
• Is the reference section provide more resources to investigate? We enrich our
scientific publication dataset accordingly (e.g., LOV4IoT-transport knowledge
repository).
The main difference between our survey and the existing ones, is that our survey
is the result of a continuous enrichment of the LOV4IoT ontology catalog since
2012 and we provide tools to support the reuse of the survey outcome (e.g., dump of
ontology code). Meanwhile, we are aware of Systematic Literature Review (SLR)
guidelines such as [58–60].
4.2 Building the Corpus of Knowledge for the Transportation
Domain
To train the dataset, we need to build a corpora of knowledge for the transportation
domain. word2vec helps in transforming texts from either scientific publications
or ontologies into vectors that can be processable by machine learning algorithms.
word2vec performs the training of the term embeddings and the process of building a
word2vec model for all identified unique words. The word2vec algorithm is based on
neural networks and builds a vocabulary from a pre-training text model and attaches
the vector representations to each word. Around 20 of the terms were not part of the
pre-trained model thus we removed those terms from the list of words. The output
of this step is the word embedding vector space representation. The genism python
library is used to implement word2vec.
Transport Ontology Dataset: We have collected 16 ontologies that can be downloaded and analyzed (as depicted in Table 1): 2 ontologies are excluded since they
Précédent

- 323/424

Suivant