Neural network
Neural networks typically comprise a large number of simple processing units
linked by weighted connections according to a specified architecture. Such networks are typically massively parallel in nature and can learn by example and then
generalize (Foody et al. 2001). In the estimates of forest biomass, a variety of
neural networks have been used. In mapping the biomass of tropical forests in
north-eastern Borneo from Landsat TM data, Foody et al. (2001) found that the
multi-layer perception (MLP) performed better than both radial basis function
(RBF) and generalized regression neural networks (GRNN). Using SPOT data in
Canada, six artificial neural network models containing between 5 and 35 neurons
all predicted forest biomass in test set with an accuracy of 60 %, R
2 = 0.80, and
RMS = 32 Mg/ha (Fraser and Li 2002).
k-Nearest Neighbor Algorithm
The k-nearest neighbor (k-NN) technique is a means to determine points that
are most similar or nearest in a covariate space (McRoberts et al. 2002, 2007;
Tomppo 1991; Tomppo and Halme, 2004). Given a set of n points, defined in real
d-dimensional space, and a query point q, the k-NN technique is simply to calculate the minimum Euclidean distances for the n points to the query point
q (Finley and Mcroberts 2008). This technique is generic, which does not specify
both the similarity or distance metric and the number of nearest neighbor sampling
units on which predictions are based. The biomass estimate in point q is calculated
by a weighted value of the biomass value of the k nearest reference points, where
the weight is determined by spectral distance. Several methods have been proposed in defining the d-dimensional space within which the nearest neighbor
search is executed (Finley and Mcroberts 2008). The methods include the
Euclidean distance or a weighted Euclidean distance (Franco-Lopez et al. 2001;
Reese et al. 2003), most similar neighbor (MSN) (Moeur and Stage 1995), and
gradient nearest neighbor (GNN) (Ohmann and Gregory 2002). The k-NN technique has been frequently applied to calculate forest attributes by combining
strategic inventory data, TM imagery, and other ancillary variables (Franco-Lopez
et al. 2001; Halme and Tomppo 2001; Katila and Tomppo 2001; McRoberts et al.
2007; Trotter et al. 1997; Fazakas et al. 1999; Ohmann and Gregory 2002; Pierce
et al. 2009). When mapping biomass from Landsat TM and inventory data in
Canada, k-NN method produces a RMSE of 59 Mg/ha compared with inventory
plots (Labrecque et al. 2006).
Regression tree model
Tree-based models, such as Random Forests (RF), establish a large number of
trees, in which different bootstrap samples of the data are used to estimate each
tree. RF constructs numerous small regression trees that vote on predictions and is
robust to over-fitting (Breiman 2001). Similar to other non-parametric approaches,
RF models are built up using training samples of biomass (such as Forest
Inventory data), satellite data, and other ancillary variables. The tree is composed
of a root node (comprised of all of the data), a set of internal nodes which are split
using a randomly selected sub-set of the predictor variables, and a set of terminal
76
X. Zhang and W. Ni-meister
Neural networks typically comprise a large number of simple processing units
linked by weighted connections according to a specified architecture. Such networks are typically massively parallel in nature and can learn by example and then
generalize (Foody et al. 2001). In the estimates of forest biomass, a variety of
neural networks have been used. In mapping the biomass of tropical forests in
north-eastern Borneo from Landsat TM data, Foody et al. (2001) found that the
multi-layer perception (MLP) performed better than both radial basis function
(RBF) and generalized regression neural networks (GRNN). Using SPOT data in
Canada, six artificial neural network models containing between 5 and 35 neurons
all predicted forest biomass in test set with an accuracy of 60 %, R
2 = 0.80, and
RMS = 32 Mg/ha (Fraser and Li 2002).
k-Nearest Neighbor Algorithm
The k-nearest neighbor (k-NN) technique is a means to determine points that
are most similar or nearest in a covariate space (McRoberts et al. 2002, 2007;
Tomppo 1991; Tomppo and Halme, 2004). Given a set of n points, defined in real
d-dimensional space, and a query point q, the k-NN technique is simply to calculate the minimum Euclidean distances for the n points to the query point
q (Finley and Mcroberts 2008). This technique is generic, which does not specify
both the similarity or distance metric and the number of nearest neighbor sampling
units on which predictions are based. The biomass estimate in point q is calculated
by a weighted value of the biomass value of the k nearest reference points, where
the weight is determined by spectral distance. Several methods have been proposed in defining the d-dimensional space within which the nearest neighbor
search is executed (Finley and Mcroberts 2008). The methods include the
Euclidean distance or a weighted Euclidean distance (Franco-Lopez et al. 2001;
Reese et al. 2003), most similar neighbor (MSN) (Moeur and Stage 1995), and
gradient nearest neighbor (GNN) (Ohmann and Gregory 2002). The k-NN technique has been frequently applied to calculate forest attributes by combining
strategic inventory data, TM imagery, and other ancillary variables (Franco-Lopez
et al. 2001; Halme and Tomppo 2001; Katila and Tomppo 2001; McRoberts et al.
2007; Trotter et al. 1997; Fazakas et al. 1999; Ohmann and Gregory 2002; Pierce
et al. 2009). When mapping biomass from Landsat TM and inventory data in
Canada, k-NN method produces a RMSE of 59 Mg/ha compared with inventory
plots (Labrecque et al. 2006).
Regression tree model
Tree-based models, such as Random Forests (RF), establish a large number of
trees, in which different bootstrap samples of the data are used to estimate each
tree. RF constructs numerous small regression trees that vote on predictions and is
robust to over-fitting (Breiman 2001). Similar to other non-parametric approaches,
RF models are built up using training samples of biomass (such as Forest
Inventory data), satellite data, and other ancillary variables. The tree is composed
of a root node (comprised of all of the data), a set of internal nodes which are split
using a randomly selected sub-set of the predictor variables, and a set of terminal
76
X. Zhang and W. Ni-meister
