10 Correlations, Hierarchies, Networks and Clustering
249
• Clustering using Random Matrix Theory (RMT) [118]; Eigenvalues help to
determine the number of clusters, and eigenvectors their composition.
• [88] proposes network-based community detection methods whose null hypothesis
is consistent with RMT results on cross-correlation matrices for financial time
series data, unlike existing community detection algorithms.
• Clustering using the p-median problem [75]; With this construction, every cluster
is a star, i.e. a tree with one central node.
On distances
At the heart of clustering algorithms is the fundamental notion of distance that can
be defined upon a proper representation of data. It is thus an obvious direction to
explore. We list below what has been proposed in the literature so far:
• Distances that try to quantify how one financial instrument provides information
about another instrument:
– Distance using Granger causality [14],
– Distance using partial correlation [72],
– Study of asynchronous, lead-lag relationships by using mutual information
instead of Pearson’s correlation coefficient [47, 124],
– The correlation matrix is normalized using the affinity transformation: the correlation between each pair of stocks is normalized according to the correlations
of each of the two stocks with all other stocks [71].
• Distances that aim at including non-linear relationships in the analysis:
– Distances using mutual information, mutual information rate, and other
information-theoretic distances [6, 9, 48, 54, 56, 124],
– The Brownian distance [154],
– Copula-based [19, 41, 94] and tail dependence [40, 84] distances.
• Distances that aim at taking into account multivariate dependence:
– Each stock is represented by a bivariate time series: its returns and traded volumes [20]; a distance is then applied to an ad hoc transform of the two time
series into a symbolic sequence,
– Each stock is represented by a multivariate time series, for example the daily
(high, low, open, close) [79]; Authors use the Escoufier’s RV coefficient (a
multivariate extension of the Pearson’s correlation coefficient).
• A distance taking into account both the correlation between returns and their
distributions [37].
• Unlike recent studies which claim that the existence of nonlinear dependence
between stock returns have effects on network characteristics, [58] documents
that “most of the apparent nonlinearity is due to univariate non-Gaussianity.
Précédent

- 257/282

Suivant