252
G. Marti et al.
• Computing Pearson correlations on a rolling window of arbitrary length,
• then independently computing a network or a clustering based on the rolling empirical correlation matrix.
Some promising avenue of research may be the use of temporal networks and temporal centrality measures [156].
Besides the shortcomings of Pearson correlation detailed above, this approach is
brittle due to its strong dependence a priori on:
• the sampling frequency (e.g., intraday, daily, weekly),
– Concerning the sampling frequency, authors in [15] notice that at intraday
frequency level some time is needed before the cluster organization emerges
completely. According to the paper, “the changes observed in the structure of
the MST and of the hierarchical tree suggest that the intrasector correlation
decreases faster than intersector correlation between pairs of stocks” when sampling frequency increases. In [17, 95], authors observe that the clusters obtained
using daily returns are similar to the ones obtained with weekly timescales, and
even to some extent to the ones using monthly returns. Most of the empirical
studies focus on daily returns and only a few explore intraday data: [15, 17, 33,
66, 80, 155]. Working with higher frequencies (e.g. at the transaction or quote
level) brings further difficulties such as coping with asynchronous data and the
Epps effect [45].
• authors study a large propthe length T of the rolling window,
– What is the right length for the rolling window? No clear-cut answer has yet been
proposed and, in most studies, its length is set somewhat arbitrarily. In [108],
authors posit that “the choice of window width is a trade-off between too noisy
and too smoothed data for small and large window widths, respectively” and
that they “have explored a large scale of different values for both parameters,
and the given values were found optimal”. What are the proper criteria for
setting the window length? The choice can be driven by the goal (e.g. time
investment horizon), by regulatory rules (e.g. computing Value-at-Risk using 1year historical data), by the stability of clusters [95], by a statistical convergence
rate [92], by economic regimes or by a trade-off of the preceding criteria.
• the number N of assets studied.
– The number of considered assets has also a significant impact on the results:
the ratio T /N drives the precision of the correlation matrix estimation and
ultimately the clustering [18, 21, 22, 92].
This dependence makes it difficult to fully understand and analyze results. Once
these ‘parameters’, i.e. the sampling frequency, T , and N , are chosen, one can study
• the dynamics of correlations:
– In [71], authors are using a sliding window of T = 22 days to measure and
monitor the eigenvalue entropy of the stock correlation matrices (estimated
G. Marti et al.
• Computing Pearson correlations on a rolling window of arbitrary length,
• then independently computing a network or a clustering based on the rolling empirical correlation matrix.
Some promising avenue of research may be the use of temporal networks and temporal centrality measures [156].
Besides the shortcomings of Pearson correlation detailed above, this approach is
brittle due to its strong dependence a priori on:
• the sampling frequency (e.g., intraday, daily, weekly),
– Concerning the sampling frequency, authors in [15] notice that at intraday
frequency level some time is needed before the cluster organization emerges
completely. According to the paper, “the changes observed in the structure of
the MST and of the hierarchical tree suggest that the intrasector correlation
decreases faster than intersector correlation between pairs of stocks” when sampling frequency increases. In [17, 95], authors observe that the clusters obtained
using daily returns are similar to the ones obtained with weekly timescales, and
even to some extent to the ones using monthly returns. Most of the empirical
studies focus on daily returns and only a few explore intraday data: [15, 17, 33,
66, 80, 155]. Working with higher frequencies (e.g. at the transaction or quote
level) brings further difficulties such as coping with asynchronous data and the
Epps effect [45].
• authors study a large propthe length T of the rolling window,
– What is the right length for the rolling window? No clear-cut answer has yet been
proposed and, in most studies, its length is set somewhat arbitrarily. In [108],
authors posit that “the choice of window width is a trade-off between too noisy
and too smoothed data for small and large window widths, respectively” and
that they “have explored a large scale of different values for both parameters,
and the given values were found optimal”. What are the proper criteria for
setting the window length? The choice can be driven by the goal (e.g. time
investment horizon), by regulatory rules (e.g. computing Value-at-Risk using 1year historical data), by the stability of clusters [95], by a statistical convergence
rate [92], by economic regimes or by a trade-off of the preceding criteria.
• the number N of assets studied.
– The number of considered assets has also a significant impact on the results:
the ratio T /N drives the precision of the correlation matrix estimation and
ultimately the clustering [18, 21, 22, 92].
This dependence makes it difficult to fully understand and analyze results. Once
these ‘parameters’, i.e. the sampling frequency, T , and N , are chosen, one can study
• the dynamics of correlations:
– In [71], authors are using a sliding window of T = 22 days to measure and
monitor the eigenvalue entropy of the stock correlation matrices (estimated
