Toward Learned Representations

The descriptors presented in the preceding sections result from explicit operator design. In all cases, a family of image transformations (gradients, binary comparisons, convolutions, histograms) is defined because it is considered particularly suitable for representing the salient information in the scene.

With the advent of Deep Learning, this paradigm was gradually replaced by approaches in which the descriptor is learned directly from the data. These systems no longer rely on a manually designed image representation; instead, a neural network is trained to produce vectors that directly maximize the ability to recognize corresponding points across different images.

More generally, a descriptor can be interpreted as a latent representation of the information contained in an image. Models such as autoencoders, RBMs (Section 5.10.2), and CNNs (Section 5.10.5) can be viewed as systems capable of compressing the observed information into a lower-dimensional space while preserving the aspects relevant to the task.

This evolution has progressively blurred the distinction between keypoint detection, description, and matching, an issue that will be examined in greater detail in the following sections.



Paolo medici
2026-10-01