Another concept that cuts across several areas of computer vision is that of the descriptor (Visual Descriptor). A descriptor is used in several contexts: to compare feature points or generate a disparity map in stereo vision, to provide a compact representation of an image region so as to speed up its detection or retrieval, and, because this compact representation preserves much of the information, to generate the feature space for classification algorithms.
Depending on the transformation undergone by the image from which the points are to be characterized, the descriptor must satisfy several invariance principles:
Before the concept of a compact descriptor was introduced, the standard method for comparing two feature points was correlation between the regions surrounding the points:
| (7.1) |
| (7.2) |
It should also be noted that comparing pixels between images is still an algorithm of type : performing these comparisons for each point nevertheless requires substantial computational effort and numerous memory accesses.
Modern approaches seek to overcome this limitation by extracting a descriptor from the point's neighborhood that is smaller than the number of pixels represented while maximizing the information it contains.
Both SIFT (Section 6.3) and SURF (Section 6.4) extract their descriptors by exploiting scale and rotation information obtained from the image (these properties can also be extracted independently, and therefore applied to any class of descriptor to make it invariant to scale and rotation). The descriptors obtained from SIFT and SURF are different versions of the same concept: the histogram of gradient orientations (Section 7.2), an example of how to compress the variability around a point into a lower-dimensional space.
None of the descriptors currently in use directly uses the image pixels as the descriptor; however, it is easy to see that a sufficiently well-distributed subset of the pixels is enough to produce an accurate description of the point.
In (RD05), a descriptor is created from the 16 pixels located along the discrete circle of radius 3.
This description can be made even more compact by converting it to the binary form of the Local Binary Pattern described below, or by using a form not constrained to the circle, as in Census or BRIEF.
Another approach is to sample the kernel space appropriately (GZS11), extracting from coordinates around the keypoint the values produced by convolutions of the original image (horizontal and vertical Sobel filters), thereby creating a descriptor with only
values.
It should be noted that, for purely computational reasons related to resource reuse, a specific descriptor extractor is often associated with each feature-point extractor.
This introduction makes it clear that describing a keypoint using fewer data while retaining sufficient descriptive power is also useful in classification. The concept of a descriptor arose from the attempt to extract local image information while preserving a substantial part of the original information. This makes it possible to perform relatively fast comparisons between points in images or to use such descriptors as features on which to train classifiers.