Statistical Modelling and Variable Selection in Climate Science
355
X =
⎛
⎜
⎜
⎜
⎝
1 x 11 x 12 · · · x 1k
1 x 21 x 22 · · · x 2k
. . .
. . .
. . .
. . .
. . .
1 x n1 x n2 · · · x nk
⎞
⎟
⎟
⎟
⎠
is a n ×(k +1) design matrix of n observations on each
of the k explanatory variables arranged in n rows and (k + 1) columns corresponding
to intercept term and k covariates, β = (β 0 , β 1 , β 2 , . . . , β k )
T is a (k + 1) × 1 vector
of regression coefficients, and ε = (ε 1 , ε 2 , . . . , ε n )
T is a n × 1 vector of random error
components or disturbance term.
Some assumptions are required in the model (3) for the implementation of statistical methods and to study the statistical properties of estimators of regression
coefficients. It is assumed that E(ε) = 0 (i.e., mean of random errors is zero),
E(εε
T
) = σ
2 I n (i.e., the random errors are identically and independently distributed
having constant variance σ
2 ), X is a full column rank non-stochastic (or fixed) matrix,
and ε ∼ N (0, σ
2 I n ) (i.e., the random errors are following a n dimensional multivariate normal distribution with mean vector 0 and covariance matrix σ
2 I n ). Note
that operator E is called as Expectation, e.g., E(ε) is called as expected value of ε,
which represents the arithmetic mean of value of all ε in the population. Another
assumption is lim
n→∞
X
T X
n
= exists, and it is a non-stochastic and nonsingular
matrix (with finite elements). Such an assumption is required to study the large sample asymptotic properties of the estimators. The explanatory variables can also be
stochastic in some cases but they are assumed to be fixed in this article.
Next, we discuss the interpretation of different regression parameters β j ( j =
0, 1, . . . , k) involved in the multiple regression model. Note that
E(y) = β 0 + β 1 X 1 + β 2 X 2 + · · · + β k X k .
(4)
We observe that the partial derivative of mean value of y, i.e., E(y) with respect to
jth explanatory variable X j , i.e.
∂ E(y)
∂ X j
= β j , j = 1, 2, . . . , k represents the expected
or average change in the response y with respect to unit change in X j , i.e., how much
the average value of y will change when the value of X j is changed by one unit. When
X j = 0, j = 1, 2, . . . , k then E(y) = β 0 that indicates the average value of y when
all the observations on all the explanatory variables are assigned zero values. Another
involved parameter is σ
2 , which measures the variation in the random error term. It
indicates the degree of variability present in the observed responses y 1 , y 2 , . . . , y n .
First, we understand that what is needed to know or find a model based on a given
set of data on study and explanatory variables. A model is a functional relationship
between y and X 1 , X 2 , . . . , X k . The functional form is unknown and described by
its parameters, which are also unknown. An attempt is made to know a possible
form of the functional relationship and then the involved parameters are needed to be
found based on a given set of data obtained from studies and experiments. Once we
come to know the values of the parameters, a functional mathematical relationship
is established between the input and output variables giving rise to a model. Though
the exact functional relationship between y and X 1 , X 2 , . . . , X k is unknown, but
Précédent

- 356/553

Suivant