E1C04 09/14/2010
14:7:43 Page 141
Consider the situation in which there are N values of x and y, referred to as x i , y i , where i ¼ 1,
2, . . . , N. We seek an mth-order polynomial based on a set of N data points of the form (x, y) in
which x and y are the independent and dependent variables, respectively. The task is to find the m þ
1 coefficients, a 0 , a 1 , . . . , a m , of the polynomial of Equation 4.31. We define the deviation between
any dependent variable y i and the polynomial as y i À y c i , where y c i is the value of the polynomial
evaluated at the data point (x i , y i ). The sum of the squares of this deviation for all values of y i , i ¼ 1,
2, . . . , N, is
D ¼
X N
i¼1
y i À y c i
À
Á 2
ð4:32Þ
The goal of the method of least-squares is to reduce D to a minimum for a given order of
polynomial.
Combining Equations 4.31 and 4.32, we write
D ¼
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
ð4:33Þ
Now the total differential of D is dependent on the m þ 1 coefficients through
dD ¼
@D
@a 0
da 0 þ
@D
@a 1
da 1 þ Á Á Á þ
@D
@a m
da m
To minimize the sum of the squares of the deviations, we want dD to be zero. This is accomplished
by setting each partial derivative equal to zero:
@D
@a 0
¼ 0 ¼
@
@a 0
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
@D
@a 1
¼ 0 ¼
@
@a 1
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
. .
.
@D
@a m
¼ 0 ¼
@
@a m
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
ð4:34Þ
This yields m þ 1 equations that are solved simultaneously to yield the unknown regression
coefficients, a 0 , a 1 , . . . , a m .
In general, the polynomial found by regression analysis does not pass through every data point
(x i , y i ) exactly, so there is some deviation, y i À y c i , between each data point and the polynomial. We
compute a standard deviation based on these differences by
s yx ¼
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
P N
i¼1 y i À y c i
À
Á 2
n
s
ð4:35Þ
where n is the degrees of freedom of the fit and n ¼ N À (m þ 1). The statistic s yx is referred to as the
standard error of the fit and is related to how closely a polynomial fits the data set.
4.6 Regression Analysis 141
14:7:43 Page 141
Consider the situation in which there are N values of x and y, referred to as x i , y i , where i ¼ 1,
2, . . . , N. We seek an mth-order polynomial based on a set of N data points of the form (x, y) in
which x and y are the independent and dependent variables, respectively. The task is to find the m þ
1 coefficients, a 0 , a 1 , . . . , a m , of the polynomial of Equation 4.31. We define the deviation between
any dependent variable y i and the polynomial as y i À y c i , where y c i is the value of the polynomial
evaluated at the data point (x i , y i ). The sum of the squares of this deviation for all values of y i , i ¼ 1,
2, . . . , N, is
D ¼
X N
i¼1
y i À y c i
À
Á 2
ð4:32Þ
The goal of the method of least-squares is to reduce D to a minimum for a given order of
polynomial.
Combining Equations 4.31 and 4.32, we write
D ¼
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
ð4:33Þ
Now the total differential of D is dependent on the m þ 1 coefficients through
dD ¼
@D
@a 0
da 0 þ
@D
@a 1
da 1 þ Á Á Á þ
@D
@a m
da m
To minimize the sum of the squares of the deviations, we want dD to be zero. This is accomplished
by setting each partial derivative equal to zero:
@D
@a 0
¼ 0 ¼
@
@a 0
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
@D
@a 1
¼ 0 ¼
@
@a 1
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
. .
.
@D
@a m
¼ 0 ¼
@
@a m
X N
i¼1
y i À a 0 þ a 1 x þ Á Á Á þ a m x
m
ð
Þ
½
2
(
)
ð4:34Þ
This yields m þ 1 equations that are solved simultaneously to yield the unknown regression
coefficients, a 0 , a 1 , . . . , a m .
In general, the polynomial found by regression analysis does not pass through every data point
(x i , y i ) exactly, so there is some deviation, y i À y c i , between each data point and the polynomial. We
compute a standard deviation based on these differences by
s yx ¼
ffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi
P N
i¼1 y i À y c i
À
Á 2
n
s
ð4:35Þ
where n is the degrees of freedom of the fit and n ¼ N À (m þ 1). The statistic s yx is referred to as the
standard error of the fit and is related to how closely a polynomial fits the data set.
4.6 Regression Analysis 141
