The reliability of the method can be checked by calculating the correlation
coefficient r with the use of (2.8) and checking how close its value is to unity. We
can use also the F-test for the null hypothesis that says that the proposed linear
relationship (5.1) fits the data set. The experimental value of the Fisher function
F can be calculated using (2.9). The corresponding experimental value F has to be
compared with the critical value F C ðk 1 ; k 2 Þ of the Fisher function at a desired
significance or confidence level (k 1 ¼ m À 1; k 2 ¼ n À m; n is the number of
observations; m = 2 is the number of parameters in the regression equation, see,
e.g., [34]). When the values n and m are fixed, the maximum of the Fisher function
coincides with the maximum of the correlation coefficient. Therefore, to find the
optimal values of parameters N and m, we have to find the maximum of the correlation coefficient for the linear dependence (5.1) or (5.3). To compare the reliability of different predictions (with different values of n), it is useful to use the ratio
F=F C 1; n À 2
ð
Þat fixed significance level. The most reliable estimation of the SIR
model parameters corresponds to the highest F=F C 1; n À 2
ð
Þvalue. This ratio was
used in [37] to find optimal links between the water animal speeds and their body
length. In further calculations, we will use the significance level 0.001; corresponding values of F C 1; n À 2
ð
Þcan be taken from [34].
The calculations of F 1j ¼ F
Ã
1 V j ; N; m
À
Á
using (4.14), linear regression coefficients
(5.3), correlation coefficient r, and F=F C 1; n À 2
ð
Þratio were performed with the
use of original MATLAB code, which allows us also to isolate the values of
parameters N and m corresponding to the maximum of r. The exact solution (4.13)–
(4.14) allows avoiding numerical solutions of differential Eqs. (4.1)–(4.3) and
significantly reduces the time spent on calculations. For large values of V j (e.g., for
the USA and the world), calculations of integral (4.14) need more computer time,
but even in these cases the calculation of one set of optimal parameters could be
performed on a regular laptop for one day.
Examples of calculations with the use of the accumulated number of cases,
confirmed in Ukraine (see Table 5.1), are presented in Table 5.2. It can be seen that
all the predictions yield very high values of the correlation coefficient (close to unit)
and high values of the F=F C 1; n À 2
ð
Þratio. The results of calculations depend on
the number of observations n (compare Predictions 4 and 5) and the data set (at
fixed values of n, compare Predictions 5 and 6). This fact reflects both the accuracy
of the approach and the changes in the epidemic behavior (for example on May 10,
2020, the national lockdown was canceled and the quarantine restrictions changed
permanently).
In SIR model, all the parameters are supposed to be constant. If the quarantine
measures and speed of isolation change or new infected persons are coming in the
country, the accuracy of the prediction reduces. It is very important to detect the
periods of time with more or less constant values of parameters, i.e., to select
different waves of the pandemic and to find optimal values of the parameters
separately for every wave. These problems will be discussed later.
Without separation of the waves, we will obtain different values of the important
epidemic characteristics. For example, the real time of the epidemic beginning t
Ã
1
5 Statistics-Based Procedure of Parameter …
35
coefficient r with the use of (2.8) and checking how close its value is to unity. We
can use also the F-test for the null hypothesis that says that the proposed linear
relationship (5.1) fits the data set. The experimental value of the Fisher function
F can be calculated using (2.9). The corresponding experimental value F has to be
compared with the critical value F C ðk 1 ; k 2 Þ of the Fisher function at a desired
significance or confidence level (k 1 ¼ m À 1; k 2 ¼ n À m; n is the number of
observations; m = 2 is the number of parameters in the regression equation, see,
e.g., [34]). When the values n and m are fixed, the maximum of the Fisher function
coincides with the maximum of the correlation coefficient. Therefore, to find the
optimal values of parameters N and m, we have to find the maximum of the correlation coefficient for the linear dependence (5.1) or (5.3). To compare the reliability of different predictions (with different values of n), it is useful to use the ratio
F=F C 1; n À 2
ð
Þat fixed significance level. The most reliable estimation of the SIR
model parameters corresponds to the highest F=F C 1; n À 2
ð
Þvalue. This ratio was
used in [37] to find optimal links between the water animal speeds and their body
length. In further calculations, we will use the significance level 0.001; corresponding values of F C 1; n À 2
ð
Þcan be taken from [34].
The calculations of F 1j ¼ F
Ã
1 V j ; N; m
À
Á
using (4.14), linear regression coefficients
(5.3), correlation coefficient r, and F=F C 1; n À 2
ð
Þratio were performed with the
use of original MATLAB code, which allows us also to isolate the values of
parameters N and m corresponding to the maximum of r. The exact solution (4.13)–
(4.14) allows avoiding numerical solutions of differential Eqs. (4.1)–(4.3) and
significantly reduces the time spent on calculations. For large values of V j (e.g., for
the USA and the world), calculations of integral (4.14) need more computer time,
but even in these cases the calculation of one set of optimal parameters could be
performed on a regular laptop for one day.
Examples of calculations with the use of the accumulated number of cases,
confirmed in Ukraine (see Table 5.1), are presented in Table 5.2. It can be seen that
all the predictions yield very high values of the correlation coefficient (close to unit)
and high values of the F=F C 1; n À 2
ð
Þratio. The results of calculations depend on
the number of observations n (compare Predictions 4 and 5) and the data set (at
fixed values of n, compare Predictions 5 and 6). This fact reflects both the accuracy
of the approach and the changes in the epidemic behavior (for example on May 10,
2020, the national lockdown was canceled and the quarantine restrictions changed
permanently).
In SIR model, all the parameters are supposed to be constant. If the quarantine
measures and speed of isolation change or new infected persons are coming in the
country, the accuracy of the prediction reduces. It is very important to detect the
periods of time with more or less constant values of parameters, i.e., to select
different waves of the pandemic and to find optimal values of the parameters
separately for every wave. These problems will be discussed later.
Without separation of the waves, we will obtain different values of the important
epidemic characteristics. For example, the real time of the epidemic beginning t
Ã
1
5 Statistics-Based Procedure of Parameter …
35
