250
G. Marti et al.
Further, strong non-stationarity in a few specific stocks may play a role. In particular, the sharp decrease of some stocks during the global financial crisis in 2008”
gives rise to apparent negative tail dependence among stocks. When constructing
unweighted stock networks, they suggest to use linear correlation “on marginally
normalized data”, that is Spearman’s rank correlation. In fact, this is similar to the
idea of splitting apart the dependence information from the distribution one as in
[37], where Spearman’s rank correlation stems from using a Euclidean distance
between the uniform margins of the underlying bivariate copula. Following previous studies, and unlike in [37], the distribution information is discarded when
constructing the network.
On other methodological aspects
Besides research contributions on algorithms and distances, other methodological
aspects have been pushed further.
• Reliability and statistical uncertainty of the methods:
– A bootstrap approach is used to estimate the statistical reliability of both hierarchical trees [92, 146] and correlation-based networks [101, 144],
– Reference [91] suggests the use of generative adversarial networks (GANs) for
sampling realistic yet artificial financial correlation matrices to compare and
test the robustness of results, an alternative approach to bootstrap methods,
– Consistency proof of clustering algorithms for recovering clusters defined by
nested block correlation matrices; Study of empirical convergence rates [92],
– Kullback-Leibler divergence is used to estimate the amount of filtered information between the sample correlation matrix and the filtered one [147],
– Cophenetic correlation is used between the original correlation distances and
the hierarchical cluster representation [115],
– Several measures between successive (in time) clusters, dendrograms, networks
are used to estimate stability of the methods, e.g. cophenetic correlation between
dendrograms in [114], adjusted Rand index (ARI) between clusters in [95],
mutual information (MI) of link co-occurrence between networks in [130].
– In [68], authors claim that clustering still cannot compete with “fundamental”
industry classifications in terms of performance due to inherent out-of-sample
instabilities, and thus propose to improve such given “fundamental” industry
classification via further clustering large sub-industries at the most granular
level; Similarly, Lopez de Prado in [86] finds that “empirical correlation matrices are unstable and backward-looking, and proposes a more sophisticated
approach to tackle this issue: A method which fits a correlation matrix constrained by a hierarchical structure encoding an economic and potentially
forward-looking hypothesis. Avellaneda, in [4], builds on similar motivations
to propose a hierarchical PCA yielding principal components which are con-
G. Marti et al.
Further, strong non-stationarity in a few specific stocks may play a role. In particular, the sharp decrease of some stocks during the global financial crisis in 2008”
gives rise to apparent negative tail dependence among stocks. When constructing
unweighted stock networks, they suggest to use linear correlation “on marginally
normalized data”, that is Spearman’s rank correlation. In fact, this is similar to the
idea of splitting apart the dependence information from the distribution one as in
[37], where Spearman’s rank correlation stems from using a Euclidean distance
between the uniform margins of the underlying bivariate copula. Following previous studies, and unlike in [37], the distribution information is discarded when
constructing the network.
On other methodological aspects
Besides research contributions on algorithms and distances, other methodological
aspects have been pushed further.
• Reliability and statistical uncertainty of the methods:
– A bootstrap approach is used to estimate the statistical reliability of both hierarchical trees [92, 146] and correlation-based networks [101, 144],
– Reference [91] suggests the use of generative adversarial networks (GANs) for
sampling realistic yet artificial financial correlation matrices to compare and
test the robustness of results, an alternative approach to bootstrap methods,
– Consistency proof of clustering algorithms for recovering clusters defined by
nested block correlation matrices; Study of empirical convergence rates [92],
– Kullback-Leibler divergence is used to estimate the amount of filtered information between the sample correlation matrix and the filtered one [147],
– Cophenetic correlation is used between the original correlation distances and
the hierarchical cluster representation [115],
– Several measures between successive (in time) clusters, dendrograms, networks
are used to estimate stability of the methods, e.g. cophenetic correlation between
dendrograms in [114], adjusted Rand index (ARI) between clusters in [95],
mutual information (MI) of link co-occurrence between networks in [130].
– In [68], authors claim that clustering still cannot compete with “fundamental”
industry classifications in terms of performance due to inherent out-of-sample
instabilities, and thus propose to improve such given “fundamental” industry
classification via further clustering large sub-industries at the most granular
level; Similarly, Lopez de Prado in [86] finds that “empirical correlation matrices are unstable and backward-looking, and proposes a more sophisticated
approach to tackle this issue: A method which fits a correlation matrix constrained by a hierarchical structure encoding an economic and potentially
forward-looking hypothesis. Avellaneda, in [4], builds on similar motivations
to propose a hierarchical PCA yielding principal components which are con-
