Toward Learned Representations

The descriptors presented in the preceding sections are the result of explicit design by the practitioner. In each case, a family of image transformations (gradients, binary comparisons, convolutions, histograms) is defined as being particularly suitable for representing the salient information in the scene.

With the advent of Deep Learning, this paradigm has gradually been replaced by approaches in which the descriptor is learned directly from data. In these systems, an image representation is no longer designed manually; instead, a neural network is trained to produce vectors that directly maximize the ability to identify corresponding points in different images.

More generally, a descriptor can be interpreted as a latent representation of the information contained in an image. Models such as autoencoders, RBMs (Section 5.10.2), and CNNs (Section 5.10.5) can be viewed as systems that compress observed information into a lower-dimensional space while preserving the aspects relevant to the task at hand.

This evolution has progressively blurred the distinction between keypoint detection, description, and matching, a topic explored further in the following sections.



Paolo medici
2026-10-06