Another concept that cuts across several areas of computer vision is the descriptor (Visual Descriptor). Descriptors are used in a variety of applications: to compare keypoints, generate disparity maps in stereo vision, and provide a compact representation of an image region to speed up its detection or search. Their compact representation, which nevertheless preserves much of the information, also makes them useful for generating the feature space in classification algorithms.
Depending on the transformations applied to the image from which the points are to be characterized, the descriptor must satisfy certain invariance properties:
Before compact descriptors were introduced, the standard way to compare two keypoints was to correlate the regions around them:
| (7.1) |
| (7.2) |
It is also worth noting that comparing pixels between images is still an algorithm: performing these comparisons for each point remains computationally expensive and requires many memory accesses.
Modern approaches aim to overcome this limitation by extracting a descriptor from the point's neighborhood that is smaller than the number of pixels represented, while maximizing the information it contains.
Both SIFT (Section 6.3) and SURF (Section 6.4) use scale and rotation information extracted from the image to produce their descriptors. (This information can also be extracted independently, and therefore applied to any class of descriptors to make them invariant to scale and rotation.) The descriptors produced by SIFT and SURF are different versions of the same concept: the gradient orientation histogram (Section 7.2), which illustrates how the variability around a point can be compressed into a low-dimensional space.
Current descriptors do not use image pixels directly as descriptor elements. However, it is easy to see that a sufficiently well-distributed subset of the pixels can still provide an accurate description of the point. In (RD05), a descriptor is formed from the 16 pixels on a discrete circle of radius 3. This representation can be made even more compact by converting it to the binary form of the Local Binary Patterns described below, or by removing the constraint that the pixels lie on a circle, as in Census or BRIEF.
Another approach is to sample the kernel space appropriately (GZS11), extracting from coordinates around the keypoint the values of convolutions of the original image (horizontal and vertical Sobel filters) to form a descriptor with just
values.
For computational efficiency and resource reuse, a specific descriptor extractor is often associated with each keypoint detector.
This introduction shows that describing a keypoint with a smaller, yet sufficiently informative, set of data is also useful in classification. Descriptors were developed to extract local image information while preserving much of the original information. This makes it possible to compare points across images relatively quickly, or to use descriptors as features for training classifiers.