E1C04 09/14/2010
14:7:39 Page 118
Chapter 4
Probability and Statistics
4.1 INTRODUCTION
Suppose we had a large box containing a population of thousands of similar-sized round bearings. To
get an idea of the size of the bearings, we might measure two dozen from these thousands. The
resulting samples of diameter values form a data set, which we then use to imply something about
the entire population of the bearings in the box, such as average size and size variation. But how
close are these values taken from our small data set to the actual average size and variation of all the
bearings in the box? If we selected a different two dozen bearings, should we expect the size values
to be exactly the same? These are questions that surround most engineering measurements. We
make some measurements that we then use to try to answer questions. What is the mean value of the
variable based on a finite number of measurements and how well does this value represent the entire
population? Do the variations found ensure that we can meet tolerances for the population as a
whole? How good are these results? These questions have answers in probability and statistics.
For a given set of measurements, we want to be able to quantify (1) a single representative value
that best characterizes the average of the measured data set; (2) a representative value that provides a
measure of the variation in the measured data set; and (3) how well the average of the data set
represents the true average value of the entire population of the measured variable. The value of item
(1) can vary with repeated data sets and so the difference between it and the true average value of the
population is a type of random error. Item (3) requires establishing an interval within which the true
average value of the population is expected to lie. This interval quantifies the probable range of this
random error and it is called a random uncertainty.
This chapter presents an introduction to the concepts of probability and statistics at a level
sufficient to provide information for a large class of engineering judgments. Such material allows for
the reduction of raw data into results.
Upon completion of this chapter, the reader will be able to
quantify the statistical characteristics of a data set as it relates to the population of the
measured variable,
explain and use probability density functions to describe the behavior of variables,
create meaningful histograms of measured data,
quantify a confidence interval about the measured mean value at a given probability,
perform regression analysis on a data set and quantify the confidence intervals for the
parameters of the resulting curve fit or response surface,
118
14:7:39 Page 118
Chapter 4
Probability and Statistics
4.1 INTRODUCTION
Suppose we had a large box containing a population of thousands of similar-sized round bearings. To
get an idea of the size of the bearings, we might measure two dozen from these thousands. The
resulting samples of diameter values form a data set, which we then use to imply something about
the entire population of the bearings in the box, such as average size and size variation. But how
close are these values taken from our small data set to the actual average size and variation of all the
bearings in the box? If we selected a different two dozen bearings, should we expect the size values
to be exactly the same? These are questions that surround most engineering measurements. We
make some measurements that we then use to try to answer questions. What is the mean value of the
variable based on a finite number of measurements and how well does this value represent the entire
population? Do the variations found ensure that we can meet tolerances for the population as a
whole? How good are these results? These questions have answers in probability and statistics.
For a given set of measurements, we want to be able to quantify (1) a single representative value
that best characterizes the average of the measured data set; (2) a representative value that provides a
measure of the variation in the measured data set; and (3) how well the average of the data set
represents the true average value of the entire population of the measured variable. The value of item
(1) can vary with repeated data sets and so the difference between it and the true average value of the
population is a type of random error. Item (3) requires establishing an interval within which the true
average value of the population is expected to lie. This interval quantifies the probable range of this
random error and it is called a random uncertainty.
This chapter presents an introduction to the concepts of probability and statistics at a level
sufficient to provide information for a large class of engineering judgments. Such material allows for
the reduction of raw data into results.
Upon completion of this chapter, the reader will be able to
quantify the statistical characteristics of a data set as it relates to the population of the
measured variable,
explain and use probability density functions to describe the behavior of variables,
create meaningful histograms of measured data,
quantify a confidence interval about the measured mean value at a given probability,
perform regression analysis on a data set and quantify the confidence intervals for the
parameters of the resulting curve fit or response surface,
118
