354
N. Balac
Once the business problem and the overall project goals are fully understood, the
project moves into the Data Understanding phase. Creating the proper dataset is the
goal of this phase. It may involve bringing together data from different sources and
of different types to be able to develop comprehensive models. The rate, quantity,
and quality of the data are carefully considered. The execution of this phase may
require reconsideration of the business understanding based on data availability,
resource limitations, and such.
The data preparation phase is frequently the most time consuming and resource
intensive phase of the process. The preprocessing and cleaning of the data undertaken in this phase can require considerable effort and should not be underestimated.
Careful, advanced planning of data collection and storage can help minimize the
effort expended in this phase.
The modeling phase can be initiated once the data has been sufficiently prepared.
However, it is typical for data preparation efforts to continue and/or be revised
based on the progress made and insights gained during the modeling process. The
modeling phase involves applying one or more data science techniques to the data
set in order to extract actionable insight.
Once models are developed (or “trained”) in the modeling phase, the evaluation
phase considers the value of the models in the context of the original business
understanding. Frequently, multiple iterations through the process are required to
arrive at a satisfactory data mining solution.
Finally, the deployment phase addresses the implementation of the models within
the organization and completes the process. This may involve multiple personnel
and expertise from a wide variety of groups in addition to the data science team.
6.6 Conclusion
Big Data is fundamentally changing the way organizations and businesses operate
and compete. Big data and IoT also share a closely knitted future to offer datadriven analysis and insight. In this chapter, we explained how to build and maintain
reliable, scalable, distributed systems with Apache Hadoop and Apache Spark. We
also discussed how to utilize Hadoop and Spark for different types of big data
analytics in IoT projects, including batch and real-time stream analysis as well as
machine learning.
References
1. https://www.ibmbigdatahub.com/infographic/extracting-business-value-4-vs-big-data
2. https://spectrum.ieee.org/tech-talk/telecom/internet/popular-internet-of-things-forecast-of-50billion-devices-by-2020-is-outdated
3. https://www.visualcapitalist.com/internet-minute-2018/
Précédent

- 359/647

Suivant