10 Correlations, Hierarchies, Networks and Clustering
247
resulting MST also provides a unique indexed hierarchy [90] which corresponds
to the one given by the dendrogram obtained using the Single Linkage Clustering
Algorithm.
10.3 Methodological Concerns and Extensions
Several papers have highlighted the potential shortcomings of the original methodology. We found many references which raise concerns, and suggest alternative methods. However, to this day, it does not seem that any alternatives have gained a broad
enough acceptance so that they are systematically adopted for empirical studies of
market correlations.
10.3.1 Concerns About the Standard Methodology
We list below the concerns that have been raised about the standard methodology
during the last 20 years:
• The clusters obtained from the MST (or equivalently, the Single Linkage Clustering
Algorithm (SLCA)) are known to be unstable (small perturbations of the input
data may cause big differences in the resulting clusters) [95].
• The clustering instability may be partly due to the algorithm (MST/Single Linkage
are known for the chaining phenomenon [26]).
• The clustering instability may be partly due to the correlation coefficient (Pearson
linear correlation) defining the distance which is known for being brittle to outliers, and, more generally, not well suited to distributions other than the Gaussian
ones [37].
• Theoretical results providing the statistical reliability of hierarchical trees and
correlation-based networks are still not available [148].
• One might expect that the higher the correlation associated to a link in a correlationbased network is, the higher the reliability of this link is. In [144], authors show
that this is not always observed empirically.
• Changes affecting specific links (and clusters) during prominent crises are of difficult interpretation due to the high level of statistical uncertainty associated
with the correlation estimation [130].
• The standard method is somewhat arbitrary: A change in the method (e.g. using
a different clustering algorithm or a different correlation coefficient) may yield a
huge change in the clustering results [81, 95]. As a consequence, it implies huge
variability in portfolio formation and perceived risk [81].
Notice that Benjamin F. King in his 1966 paper [74] (the first paper, to the best of
our knowledge, about clustering stocks based on their historical returns; apparently
unknown to Mantegna and his colleagues who reinvented a similar method) adds
247
resulting MST also provides a unique indexed hierarchy [90] which corresponds
to the one given by the dendrogram obtained using the Single Linkage Clustering
Algorithm.
10.3 Methodological Concerns and Extensions
Several papers have highlighted the potential shortcomings of the original methodology. We found many references which raise concerns, and suggest alternative methods. However, to this day, it does not seem that any alternatives have gained a broad
enough acceptance so that they are systematically adopted for empirical studies of
market correlations.
10.3.1 Concerns About the Standard Methodology
We list below the concerns that have been raised about the standard methodology
during the last 20 years:
• The clusters obtained from the MST (or equivalently, the Single Linkage Clustering
Algorithm (SLCA)) are known to be unstable (small perturbations of the input
data may cause big differences in the resulting clusters) [95].
• The clustering instability may be partly due to the algorithm (MST/Single Linkage
are known for the chaining phenomenon [26]).
• The clustering instability may be partly due to the correlation coefficient (Pearson
linear correlation) defining the distance which is known for being brittle to outliers, and, more generally, not well suited to distributions other than the Gaussian
ones [37].
• Theoretical results providing the statistical reliability of hierarchical trees and
correlation-based networks are still not available [148].
• One might expect that the higher the correlation associated to a link in a correlationbased network is, the higher the reliability of this link is. In [144], authors show
that this is not always observed empirically.
• Changes affecting specific links (and clusters) during prominent crises are of difficult interpretation due to the high level of statistical uncertainty associated
with the correlation estimation [130].
• The standard method is somewhat arbitrary: A change in the method (e.g. using
a different clustering algorithm or a different correlation coefficient) may yield a
huge change in the clustering results [81, 95]. As a consequence, it implies huge
variability in portfolio formation and perceived risk [81].
Notice that Benjamin F. King in his 1966 paper [74] (the first paper, to the best of
our knowledge, about clustering stocks based on their historical returns; apparently
unknown to Mantegna and his colleagues who reinvented a similar method) adds
