Chapter VI
Machine Learning
84
VI.8. Model developed for rubble mound breakwater
The cleaned database, containing 8091 entries with different types of structures, was reduced by
selecting only the structures most similar to the rubble mound breakwater implemented at
Marina and Port of Algiers. Furthermore, the data was filtered based on their reliability factor,
with tests having reliability levels 1, 2, and 3 being chosen.
In order to improve the prediction performance of wave overtopping and achieve a good score,
the least influential features are eliminated. The hydraulic parameters of deep waters were not
considered to avoid redundancy, as they are strongly related to the parameters at the base of the
structure. However, tests corresponding to overtopping with í µí± = 0 í µí±
3 /s/m were included. The
resulting database now have 848 entries, and the number of parameters has been reduced from
31 to 13 parameters. The Figure VI-12 provides an overview of the database used during the
model development.
Figure VI-12 Overview of the Rubble mound breakwater database.
Firstly, this database was divided into two parts, with 80% of the data used for training and 20%
for testing. Then, the features were subjected to a polynomial transformation of degree 4 (Amara,
L, Chalal, Y. (2022)), which generated a new feature vector.
Various methods, such as polynomial regression, support vector regression, neural network (ML
Regressor), and XGBoost, were tested on the database. Unfortunately, the obtained results were
not satisfactory.
Afterwards, the XGBoost model using scikit-learn was applied with the following parameters.
• max_depth: 6. This parameter controls the maximum depth of a tree. A higher value can
result in a more complex model, which may lead to overfitting.
• min_child_weight: 5. This parameter specifies the minimum sum of instance weight
(hessian) needed in a child. It is used to control over-fitting. A higher value makes the
algorithm more conservative.
• learning_rate: 0.05. This parameter controls the step size shrinkage during each boosting
iteration. A lower value makes the model more robust but requires more iterations to
converge.
Machine Learning
84
VI.8. Model developed for rubble mound breakwater
The cleaned database, containing 8091 entries with different types of structures, was reduced by
selecting only the structures most similar to the rubble mound breakwater implemented at
Marina and Port of Algiers. Furthermore, the data was filtered based on their reliability factor,
with tests having reliability levels 1, 2, and 3 being chosen.
In order to improve the prediction performance of wave overtopping and achieve a good score,
the least influential features are eliminated. The hydraulic parameters of deep waters were not
considered to avoid redundancy, as they are strongly related to the parameters at the base of the
structure. However, tests corresponding to overtopping with í µí± = 0 í µí±
3 /s/m were included. The
resulting database now have 848 entries, and the number of parameters has been reduced from
31 to 13 parameters. The Figure VI-12 provides an overview of the database used during the
model development.
Figure VI-12 Overview of the Rubble mound breakwater database.
Firstly, this database was divided into two parts, with 80% of the data used for training and 20%
for testing. Then, the features were subjected to a polynomial transformation of degree 4 (Amara,
L, Chalal, Y. (2022)), which generated a new feature vector.
Various methods, such as polynomial regression, support vector regression, neural network (ML
Regressor), and XGBoost, were tested on the database. Unfortunately, the obtained results were
not satisfactory.
Afterwards, the XGBoost model using scikit-learn was applied with the following parameters.
• max_depth: 6. This parameter controls the maximum depth of a tree. A higher value can
result in a more complex model, which may lead to overfitting.
• min_child_weight: 5. This parameter specifies the minimum sum of instance weight
(hessian) needed in a child. It is used to control over-fitting. A higher value makes the
algorithm more conservative.
• learning_rate: 0.05. This parameter controls the step size shrinkage during each boosting
iteration. A lower value makes the model more robust but requires more iterations to
converge.
