Descriptors

Another concept that cuts across several areas of computer vision is the descriptor (Visual Descriptor). Descriptors are used in a variety of applications: to compare keypoints, generate disparity maps in stereo vision, and provide a compact representation of an image region to speed up its detection or search. Their compact representation, which nevertheless preserves much of the information, also makes them useful for generating the feature space in classification algorithms.

Depending on the transformations applied to the image from which the points are to be characterized, the descriptor must satisfy certain invariance properties:

translation
This is the simplest transformation and is handled automatically by the keypoint detector;
scale
This transformation is also normally handled by the keypoint detector;
brightness
Images may undergo changes in brightness;
rotation
Images may depict the same scene rotated;
perspective
Changes in perspective produce complex deformations of the observed scene.

Before compact descriptors were introduced, the standard way to compare two keypoints was to correlate the regions around them:

\begin{displaymath}
d(\mathbf{p}_1,\mathbf{p}_2)=\sum_{\boldsymbol\delta \in \O...
...\bar {I_1})(I_2(\mathbf{p}_2 + \boldsymbol\delta) - \bar{I_2})
\end{displaymath} (7.1)

where $\Omega$ is a fixed-size window centered on the point in each of the two images, and $\bar{I}_n$ is the mean image intensity within the window $\Omega$. $w_{\delta}$ is an optional weight (for example, a Gaussian) that assigns different weights to pixels according to their distance from the point. Correlation is invariant to changes in brightness but computationally expensive. In this case, the descriptor is simply the image region around the detected point (Mor80). A similar approach, which is not invariant to brightness but is more computationally efficient, is SAD (Sum of Absolute Differences):
\begin{displaymath}
d(\mathbf{p}_1,\mathbf{p}_2)=\sum_{\boldsymbol\delta \in \O...
...oldsymbol\delta) - I_2(\mathbf{p}_2 + \boldsymbol\delta) \vert
\end{displaymath} (7.2)

To make SAD invariant to brightness, comparisons are usually performed not on the original image but on its horizontal and vertical derivatives. This simple idea can be generalized by performing comparisons not on the original image, but between one or more images obtained using different kernels. These kernels provide the descriptor with various degrees of invariance.

It is also worth noting that comparing pixels between images is still an $O(n^2)$ algorithm: performing these comparisons for each point remains computationally expensive and requires many memory accesses. Modern approaches aim to overcome this limitation by extracting a descriptor from the point's neighborhood that is smaller than the number of pixels represented, while maximizing the information it contains.

Both SIFT (Section 6.3) and SURF (Section 6.4) use scale and rotation information extracted from the image to produce their descriptors. (This information can also be extracted independently, and therefore applied to any class of descriptors to make them invariant to scale and rotation.) The descriptors produced by SIFT and SURF are different versions of the same concept: the gradient orientation histogram (Section 7.2), which illustrates how the variability around a point can be compressed into a low-dimensional space.

Current descriptors do not use image pixels directly as descriptor elements. However, it is easy to see that a sufficiently well-distributed subset of the pixels can still provide an accurate description of the point. In (RD05), a descriptor is formed from the 16 pixels on a discrete circle of radius 3. This representation can be made even more compact by converting it to the binary form of the Local Binary Patterns described below, or by removing the constraint that the pixels lie on a circle, as in Census or BRIEF. Another approach is to sample the kernel space appropriately (GZS11), extracting from $m$ coordinates around the keypoint the values of convolutions of the original image (horizontal and vertical Sobel filters) to form a descriptor with just $2m$ values.

For computational efficiency and resource reuse, a specific descriptor extractor is often associated with each keypoint detector.

This introduction shows that describing a keypoint with a smaller, yet sufficiently informative, set of data is also useful in classification. Descriptors were developed to extract local image information while preserving much of the original information. This makes it possible to compare points across images relatively quickly, or to use descriptors as features for training classifiers.



Subsections
Paolo medici
2026-10-06