Chapter VI
Machine Learning
72
Table VI-2 Linear regression formula
VI.3.2. Polynomial regression
Polynomial regression extends the concept of linear regression and allows for the modeling of
nonlinear relationships between variables. In polynomial regression, the relationship between the
dependent variable and the independent variables is expressed as an nth degree polynomial
equation. Although the equation itself is nonlinear, polynomial regression can be treated as a
linear problem by transforming the features appropriately.
Table VI-3 presents an example of a second-order polynomial regression model with two
features, illustrating how polynomial regression can be applied to capture complex relationships
between variables.
Table VI-3 Example of converting a polynomial problem to a linear regression
Example: 2 features with a second order polynomial
í µí±¦ = í µí¼ 0 + í µí¼ 1 í µí±¥ 1 + í µí¼ 2 í µí±¥ 2 + í µí¼ 3 í µí±¥ 1 í µí±¥ 2 + í µí¼ 4 í µí±¥ 1
2
+ í µí¼ 5 í µí±¥ 2
2
⟹ í µí±¦ = í µí¼ 0 + í µí¼ 1 í µí±¥ 1 + í µí¼ 2 í µí±¥ 2 + í µí¼ 3 í µí±¥ 3 + í µí¼ 4 í µí±¥ 4 + í µí¼ 5 í µí±¥ 5
Where:
í µí±¥ 3 = í µí±¥ 1 í µí±¥ 2 ; í µí±¥ 4 = í µí±¥ 1
2
; í µí±¥ 5 = í µí±¥ 2
2
;
So, a second polynomial problem of two features become a linear problem of five features.
VI.3.3. Decision Trees
The construction of a Decision Tree involves the creation of a tree-like model to predict and
classify data. It begins with a single node known as the root node, which represents the entire
dataset (Figure VI-2). The algorithm then evaluates the available features and selects the most
informative one to split the data into two or more branches. This splitting process is repeated
recursively for each resulting branch, creating additional nodes and branches until a stopping
criterion is met. The stopping criterion could be reaching a maximum depth, achieving a
minimum number of samples in a leaf node, or reaching a specific level of impurity reduction.
At each node, the algorithm makes decisions based on the selected features, aiming to maximize
the separation of classes or minimize impurity. The construction process continues until all the
data is accurately classified or until the stopping criterion is satisfied.
Machine Learning
72
Table VI-2 Linear regression formula
VI.3.2. Polynomial regression
Polynomial regression extends the concept of linear regression and allows for the modeling of
nonlinear relationships between variables. In polynomial regression, the relationship between the
dependent variable and the independent variables is expressed as an nth degree polynomial
equation. Although the equation itself is nonlinear, polynomial regression can be treated as a
linear problem by transforming the features appropriately.
Table VI-3 presents an example of a second-order polynomial regression model with two
features, illustrating how polynomial regression can be applied to capture complex relationships
between variables.
Table VI-3 Example of converting a polynomial problem to a linear regression
Example: 2 features with a second order polynomial
í µí±¦ = í µí¼ 0 + í µí¼ 1 í µí±¥ 1 + í µí¼ 2 í µí±¥ 2 + í µí¼ 3 í µí±¥ 1 í µí±¥ 2 + í µí¼ 4 í µí±¥ 1
2
+ í µí¼ 5 í µí±¥ 2
2
⟹ í µí±¦ = í µí¼ 0 + í µí¼ 1 í µí±¥ 1 + í µí¼ 2 í µí±¥ 2 + í µí¼ 3 í µí±¥ 3 + í µí¼ 4 í µí±¥ 4 + í µí¼ 5 í µí±¥ 5
Where:
í µí±¥ 3 = í µí±¥ 1 í µí±¥ 2 ; í µí±¥ 4 = í µí±¥ 1
2
; í µí±¥ 5 = í µí±¥ 2
2
;
So, a second polynomial problem of two features become a linear problem of five features.
VI.3.3. Decision Trees
The construction of a Decision Tree involves the creation of a tree-like model to predict and
classify data. It begins with a single node known as the root node, which represents the entire
dataset (Figure VI-2). The algorithm then evaluates the available features and selects the most
informative one to split the data into two or more branches. This splitting process is repeated
recursively for each resulting branch, creating additional nodes and branches until a stopping
criterion is met. The stopping criterion could be reaching a maximum depth, achieving a
minimum number of samples in a leaf node, or reaching a specific level of impurity reduction.
At each node, the algorithm makes decisions based on the selected features, aiming to maximize
the separation of classes or minimize impurity. The construction process continues until all the
data is accurately classified or until the stopping criterion is satisfied.
