When only summary statistics are available, one must be acutely aware
of how the data were collected. Hilborn and Mangel (1997) provide an
example of using summary data while assuming a normal distribution.
In their example, when the raw data were correctly fit to a negative
binomial distribution, a very large change in management action was
recommended.
Additionally, calculating descriptive statistics may provide quantitative
parameters to describe the observed patterns in plots of data. It may be
useful to know how the sample mean and variance are related. Some inference of the underlying probability distribution of the data can be gleaned
by calculating the coefficient of variation (CV), which is sometimes a useful
measurement for displaying the variance-to-mean ratio. If the variance
increases as the mean increases but the CV does not, this can indicate
that the distribution may be log-normal rather than normal. The CV may
become unstable when data are near the detection limit.
10.4.2 Outliers
One benefit of plotting the data is that outliers may become more evident.
In descriptive statistics, the univariate descriptions of mean, mode, standard
deviation, etc., may also help identify outliers. Once identified, outliers can
be removed from the data set. However this must be done with caution,
because outliers sometimes indicate real properties of the data, such as distributional asymmetry, or can be an “exception to the rule,” which, upon
further investigation, leads to new information or deeper understanding of
the phenomena under study. It is always best if removal of the outlier can
be justified on a basis other than its simply being an outlier. The existence
of good metadata can greatly assist in this decision process. While there are
quantitative methods to deal with outliers, there is always some subjectivity involved with the process (Little and Rubin 1983). In the multivariate
case, some care must be taken to identify outliers because the outliers may
not be identified with univariate methods over a series of variables. Robust
methods to deal with outliers may use other statistics besides the mean and
variance, such as the median and median deviation (Cressie 1993). In
any case, managers should employ professional statisticians in developing
models and particularly in the identification and removal of outliers in
multivariate analysis.
10.4.3 Missing Data and Imputation
If outliers were removed or observations were simply not recorded, there
will be missing data. There are several approaches to dealing with missing
data: delete the record, with some loss in the quantity of information; delete
the point while preserving the remaining information in the record; use a
196
David Hohler et al.
of how the data were collected. Hilborn and Mangel (1997) provide an
example of using summary data while assuming a normal distribution.
In their example, when the raw data were correctly fit to a negative
binomial distribution, a very large change in management action was
recommended.
Additionally, calculating descriptive statistics may provide quantitative
parameters to describe the observed patterns in plots of data. It may be
useful to know how the sample mean and variance are related. Some inference of the underlying probability distribution of the data can be gleaned
by calculating the coefficient of variation (CV), which is sometimes a useful
measurement for displaying the variance-to-mean ratio. If the variance
increases as the mean increases but the CV does not, this can indicate
that the distribution may be log-normal rather than normal. The CV may
become unstable when data are near the detection limit.
10.4.2 Outliers
One benefit of plotting the data is that outliers may become more evident.
In descriptive statistics, the univariate descriptions of mean, mode, standard
deviation, etc., may also help identify outliers. Once identified, outliers can
be removed from the data set. However this must be done with caution,
because outliers sometimes indicate real properties of the data, such as distributional asymmetry, or can be an “exception to the rule,” which, upon
further investigation, leads to new information or deeper understanding of
the phenomena under study. It is always best if removal of the outlier can
be justified on a basis other than its simply being an outlier. The existence
of good metadata can greatly assist in this decision process. While there are
quantitative methods to deal with outliers, there is always some subjectivity involved with the process (Little and Rubin 1983). In the multivariate
case, some care must be taken to identify outliers because the outliers may
not be identified with univariate methods over a series of variables. Robust
methods to deal with outliers may use other statistics besides the mean and
variance, such as the median and median deviation (Cressie 1993). In
any case, managers should employ professional statisticians in developing
models and particularly in the identification and removal of outliers in
multivariate analysis.
10.4.3 Missing Data and Imputation
If outliers were removed or observations were simply not recorded, there
will be missing data. There are several approaches to dealing with missing
data: delete the record, with some loss in the quantity of information; delete
the point while preserving the remaining information in the record; use a
196
David Hohler et al.
