352
Shalabh and S. S. Dhar
and outcomes of the phenomenon and events. This step provided a giant leap in
the knowledge discovery in understanding the nature, and hence, the climate. For
example, the cutting of trees is a factor, which significantly affects the rain but the size
of tyre of a vehicle cannot influence the rain. Similarly, an excellent urban planning
can change the average atmospheric temperature of a city; more plantation of trees and
vegetation can decrease the average atmospheric temperature; how the atmospheric
pressure, temperature, wind speed, etc. can affect or forecast the rainfall or climate
change, etc. The question here is that how do we decide over such conclusions.
Just observing the phenomenon and without the support of any quantified scientific
analysis will not provide believable information. The role of Statistical Science and
its tools become crucial in such a decision-making process and provides a scientific
basis to have belief on such conclusive pieces of evidence.
As per Berliner (2003), there are three stages to understand climate change and
related uncertainties—understanding the natural climate variability, experimentation
and using primary information sources through observations. Various attempts have
been made in the literature to construct climate models using multiple statistical
strategies and an important aspect in the climatic research is its statistical modeling.
Suppose that the outcome of any experiment or phenomenon depends on input variables. The modeling aims at determining the mathematical functional relationship
between the output and input variables. Various statistical techniques are available,
which provide statistical modeling from different perspectives, viz., parametric, nonparametric, Bayesian, frequentist, etc. for linear as well as nonlinear models. Various
methodologies have been proposed in the statistical literature to obtain a model based
on a given set of data on input and output variables. Among them, linear regression
analysis is a popular technique, which uses the data on input and output variables to
find a linear model.
The first step in any modelling is to identify the variables, which are causing or
affecting the outcome of a phenomenon. Usually, there are many variables, which
affect the outcome in climate science. Some of those variables are more important
in the sense that they explain the variability in the outcome better than the other
variables. Also, whenever an experimenter tries to find a model, the aim is to find out
a model, which is as close as possible to the outcomes in the real world. Due to this
overenthusiasm, the experimenter considers a large number of variables and collects
observations on them. Obviously, considering a large number of variables results
in more cost of experimentation, time, labour and finally more complications in
computations. Such a problem can be avoided by considering only those “important”
variables, which are contributing significantly in understanding and explaining the
variation in the model as well as avoiding less “important” variables. In the context of
linear regression modelling, several methodologies like forward selection, backward
elimination, stepwise regression, etc. are popularly used, but they are useful when the
number of variables affecting the outcome is not too large. Such classical approaches
do not work satisfactorily when the number of variables affecting the variables is
large. The approach of LASSO (least absolute shrinkage and selection operator),
proposed by Tibshirani (1996) (see also Hastie et al. 2009, 2015) attempted to solve
the issue of selecting the subset of “important” variables from a pool of all possible
Précédent

- 353/553

Suivant