256
F. Firouzi et al.
Fig. 5.9 Clustering of engine data
Business & Data
Understanding
Data Preparation
Modeling &
Evaluation
Deployment
Fig. 5.10 Machine learning flow
5.1.6 Machine Learning Flow
5.1.6.1 Overall Flow of Machine Learning Projects
The most common methodology for machine learning projects consists of the
following phases (see Fig. 5.10). It should be noted that in the data mining context,
this methodology is known as cross-industry standard process (CRISP).
• Business and Data Understanding: In this phase, we need to define the scope
of the project; understand the problem statement and pain points; study those
factors which might be able to impact the project; construct, gather, and collect
data from several sources (e.g., sensor readouts from IoT devices); and identify
metrics and key performance indicators (KPIs) for measuring success.
• Data Preparation: In this phase, we need to prepare the data for machine learning
algorithms. This phase includes (but not limited to) formatting data according to
our machine learning algorithms, handling missing values, handling categorical
variables, data normalization, and preparing training and test datasets.
• Modeling and Evaluation: In this phase, we build several machine learning
models, evaluate the performance of each of which, and finally select the best
F. Firouzi et al.
Fig. 5.9 Clustering of engine data
Business & Data
Understanding
Data Preparation
Modeling &
Evaluation
Deployment
Fig. 5.10 Machine learning flow
5.1.6 Machine Learning Flow
5.1.6.1 Overall Flow of Machine Learning Projects
The most common methodology for machine learning projects consists of the
following phases (see Fig. 5.10). It should be noted that in the data mining context,
this methodology is known as cross-industry standard process (CRISP).
• Business and Data Understanding: In this phase, we need to define the scope
of the project; understand the problem statement and pain points; study those
factors which might be able to impact the project; construct, gather, and collect
data from several sources (e.g., sensor readouts from IoT devices); and identify
metrics and key performance indicators (KPIs) for measuring success.
• Data Preparation: In this phase, we need to prepare the data for machine learning
algorithms. This phase includes (but not limited to) formatting data according to
our machine learning algorithms, handling missing values, handling categorical
variables, data normalization, and preparing training and test datasets.
• Modeling and Evaluation: In this phase, we build several machine learning
models, evaluate the performance of each of which, and finally select the best
