92
H. K. Palo
Table 2
Various machine learning algorithms and their attributes
ML algorithms
Attributes
Pros.
Cons.
Decision tree
• Uses a greedy top-down tree
approach
• The information gain is estimated at
each node
• The highest score of attributes is
chosen
• A comprehensive analysis
• Easy to use, simple to understand and
interpret
• Assign specific values to each
problem reduces uncertainty, clears
up the ambiguity, A white-box model
• Unstable
• Relatively inaccurate
• The information gain is biased
• Calculations can get very complex,
particularly if many values are
uncertain
Random forest
• Based on the bagging algorithm
• Uses an Ensemble learning technique
• Longer Training Period
• Complex
• It can handle large data
• Robust and scalable
• Solve both classifications as well as
regression problems
• Works well with both categorical and
continuous variables
• It can automatically handle missing
values
• No feature scaling required
• Handles non-linear parameters
efficiently
• Robust to outliers
• Less impacted by noise
K-nearest neighbors (KNN)
• Instance-based learning approach
• The most frequent class is
determined based on K-nearest
neighbors of an input feature
• Accurate and insensitivity to outliers
• Suitable for both numerical and
nominal features
• Noise intolerance
• Memory insensitive
• Curse of dimensionality
• Subject to boundary complexity
(continued)
H. K. Palo
Table 2
Various machine learning algorithms and their attributes
ML algorithms
Attributes
Pros.
Cons.
Decision tree
• Uses a greedy top-down tree
approach
• The information gain is estimated at
each node
• The highest score of attributes is
chosen
• A comprehensive analysis
• Easy to use, simple to understand and
interpret
• Assign specific values to each
problem reduces uncertainty, clears
up the ambiguity, A white-box model
• Unstable
• Relatively inaccurate
• The information gain is biased
• Calculations can get very complex,
particularly if many values are
uncertain
Random forest
• Based on the bagging algorithm
• Uses an Ensemble learning technique
• Longer Training Period
• Complex
• It can handle large data
• Robust and scalable
• Solve both classifications as well as
regression problems
• Works well with both categorical and
continuous variables
• It can automatically handle missing
values
• No feature scaling required
• Handles non-linear parameters
efficiently
• Robust to outliers
• Less impacted by noise
K-nearest neighbors (KNN)
• Instance-based learning approach
• The most frequent class is
determined based on K-nearest
neighbors of an input feature
• Accurate and insensitivity to outliers
• Suitable for both numerical and
nominal features
• Noise intolerance
• Memory insensitive
• Curse of dimensionality
• Subject to boundary complexity
(continued)
