patches. When a patch is located on a low contrast area, all close patches of a
given size are similar. When the patch is located on a line or edge, close patches
along the edge are similar too. This similarity is measured with the sum of
squared differences of multiple patches being compared.
• Affine variants of the previous methods can be used to identify similar regions
between images with scale change, rotation and shearing. Harris affine and
Hessian affine are the most common.
7.3.2.2 Construction of the Descriptor
Once interest points are detected, they must be made distinctive. To do so, one
approach is to build a vector which depends on a small image patch near the interest
point. There are two common methods used to compute a descriptor: SIFT and
SURF. The SIFT descriptor is based on orientation histograms. The SURF
(Speed-Up Robust Features) [30] descriptor is based on the sum of the Haar wavelet
response around the point of interest.
In the SIFT algorithm, the calculation of the descriptor is similar to the last step
of the detection process, but it is more computationally intensive. The descriptor is
computed from an image area around the interest point, which is first transformed
according to the scale and orientation computed during the detection phase. This
ensures that the content of the descriptor is not sensitive to the image scaling and
rotation. Thus, the descriptor only depends on the orientation and the amplitude of
the gradient in multiple neighborhood areas of the point of interest. In order to
create a descriptor having the aforementioned properties a common way to proceed
is to determine the orientations in different 4 Â 4 pixels patches in the 16 Â 16
neighborhood of the interest point. Nevertheless, other patch sizes may be used.
The Fig. 7.14 shows an example of gradient in each patch around an interest point.
An histogram of orientations is then created for each patch. The final descriptor is
composed by a vector containing values from all histograms.
The final size of the descriptor depends on the choice of patch size and the
number of orientations classes in the histogram. For example, for a patch size of
4 Â 4 pixels and a number of orientation classes of 8, the histogram created will
have a size of 36 like in the Fig. 7.14.
In this example, the values in each histogram of the 16 patches are organized
according to 8 angles (0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°) for a final size
vector of 128 values. Then the vector is normalized in sum unit to obtain contrast
invariance. To be sure that the vector of the interest point will be invariant in front
of the local affine transformations, each value is weighted in the histogram. The
method allows having a descriptor more robust and unique for each interest point
on the image.
As in the detection section, several other alternatives at the SIFT descriptor are
available since some years. These alternatives are presented here as a non-exhaustive
list:
200
A. Verguet et al.
given size are similar. When the patch is located on a line or edge, close patches
along the edge are similar too. This similarity is measured with the sum of
squared differences of multiple patches being compared.
• Affine variants of the previous methods can be used to identify similar regions
between images with scale change, rotation and shearing. Harris affine and
Hessian affine are the most common.
7.3.2.2 Construction of the Descriptor
Once interest points are detected, they must be made distinctive. To do so, one
approach is to build a vector which depends on a small image patch near the interest
point. There are two common methods used to compute a descriptor: SIFT and
SURF. The SIFT descriptor is based on orientation histograms. The SURF
(Speed-Up Robust Features) [30] descriptor is based on the sum of the Haar wavelet
response around the point of interest.
In the SIFT algorithm, the calculation of the descriptor is similar to the last step
of the detection process, but it is more computationally intensive. The descriptor is
computed from an image area around the interest point, which is first transformed
according to the scale and orientation computed during the detection phase. This
ensures that the content of the descriptor is not sensitive to the image scaling and
rotation. Thus, the descriptor only depends on the orientation and the amplitude of
the gradient in multiple neighborhood areas of the point of interest. In order to
create a descriptor having the aforementioned properties a common way to proceed
is to determine the orientations in different 4 Â 4 pixels patches in the 16 Â 16
neighborhood of the interest point. Nevertheless, other patch sizes may be used.
The Fig. 7.14 shows an example of gradient in each patch around an interest point.
An histogram of orientations is then created for each patch. The final descriptor is
composed by a vector containing values from all histograms.
The final size of the descriptor depends on the choice of patch size and the
number of orientations classes in the histogram. For example, for a patch size of
4 Â 4 pixels and a number of orientation classes of 8, the histogram created will
have a size of 36 like in the Fig. 7.14.
In this example, the values in each histogram of the 16 patches are organized
according to 8 angles (0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°) for a final size
vector of 128 values. Then the vector is normalized in sum unit to obtain contrast
invariance. To be sure that the vector of the interest point will be invariant in front
of the local affine transformations, each value is weighted in the histogram. The
method allows having a descriptor more robust and unique for each interest point
on the image.
As in the detection section, several other alternatives at the SIFT descriptor are
available since some years. These alternatives are presented here as a non-exhaustive
list:
200
A. Verguet et al.
