356
Shalabh and S. S. Dhar
over specific ranges of the explanatory variables, the linear regression model is an
adequate approximation to the true unknown functional relationship. Multiple linear
regression models are often used as empirical models or approximating functions to
describe the relationship between the input and output variables.
Finding a complete model involves various steps and issues. The first step is to
assume the possible functional relationship between y and X 1 , X 2 , . . . , X k . Next
step is to use an appropriate statistical estimation technique to find out the values
of involved parameters as point estimates and/or interval estimates. This is followed
by the test of hypothesis for the estimated parameters. Then the goodness of fitted
model is checked. In between, some other issues crop up in this process, e.g., how to
know if the model obtained is good or not, the variables X 1 , X 2 , . . . , X k are relevant
or not, the choice of k, i.e., the number of relevant variables, etc. Addressing such
issues adequately is also a part of modelling. For example, changing the value of
k will change the values of parameters, and consequently, the decision on the test
of hypothesis about the parameters, goodness of fit, etc. may also be altered. This
will result into a new model. Moreover, all the aspects are interrelated, and usually,
a good model cannot be obtained in a single shot. Instead, it is a recursive process
and continues until the experimenter is satisfied with the model in the sense that the
model is representing the real world phenomenon as close as possible.
2.2 Point Estimation of Parameters
The problem of knowing the unknown parameters based on a given set of data can
then be formulated as finding a “good” value for the coefficient vector β, which is the
unknown parameter to be estimated from the given data. There are various methods
of estimation available in the literature, which gives rise to different forms of the
estimators. We use the ordinary least square method for finding out the value of the
parameters, i.e., parameter estimation. This provides the values of the parameters as
a point estimate, i.e., a single value. A general procedure for the estimation of the
regression coefficient vector is to minimize the random errors as
n
i=1
f (ε i ) =
n
i=1
f (y i − β 0 − x i1 β 1 − x i2 β 2 − · · · − x ik β k )
(5)
for a suitably chosen function f. Another well known choice of f is f (ε) = |ε|
leading to the least absolute deviation estimator. We consider the principle of least
square, which is associated with f (ε) = ε
2
. We minimize the sum of squared errors
ε i ’s in Y = Xβ + , i.e.,
S(β) =
n
i=1
ε
2
i = ε
T
ε = (y − Xβ)
T
(y − Xβ)
(6)
Shalabh and S. S. Dhar
over specific ranges of the explanatory variables, the linear regression model is an
adequate approximation to the true unknown functional relationship. Multiple linear
regression models are often used as empirical models or approximating functions to
describe the relationship between the input and output variables.
Finding a complete model involves various steps and issues. The first step is to
assume the possible functional relationship between y and X 1 , X 2 , . . . , X k . Next
step is to use an appropriate statistical estimation technique to find out the values
of involved parameters as point estimates and/or interval estimates. This is followed
by the test of hypothesis for the estimated parameters. Then the goodness of fitted
model is checked. In between, some other issues crop up in this process, e.g., how to
know if the model obtained is good or not, the variables X 1 , X 2 , . . . , X k are relevant
or not, the choice of k, i.e., the number of relevant variables, etc. Addressing such
issues adequately is also a part of modelling. For example, changing the value of
k will change the values of parameters, and consequently, the decision on the test
of hypothesis about the parameters, goodness of fit, etc. may also be altered. This
will result into a new model. Moreover, all the aspects are interrelated, and usually,
a good model cannot be obtained in a single shot. Instead, it is a recursive process
and continues until the experimenter is satisfied with the model in the sense that the
model is representing the real world phenomenon as close as possible.
2.2 Point Estimation of Parameters
The problem of knowing the unknown parameters based on a given set of data can
then be formulated as finding a “good” value for the coefficient vector β, which is the
unknown parameter to be estimated from the given data. There are various methods
of estimation available in the literature, which gives rise to different forms of the
estimators. We use the ordinary least square method for finding out the value of the
parameters, i.e., parameter estimation. This provides the values of the parameters as
a point estimate, i.e., a single value. A general procedure for the estimation of the
regression coefficient vector is to minimize the random errors as
n
i=1
f (ε i ) =
n
i=1
f (y i − β 0 − x i1 β 1 − x i2 β 2 − · · · − x ik β k )
(5)
for a suitably chosen function f. Another well known choice of f is f (ε) = |ε|
leading to the least absolute deviation estimator. We consider the principle of least
square, which is associated with f (ε) = ε
2
. We minimize the sum of squared errors
ε i ’s in Y = Xβ + , i.e.,
S(β) =
n
i=1
ε
2
i = ε
T
ε = (y − Xβ)
T
(y − Xβ)
(6)
