The methods described in the preceding chapters typically involve three distinct stages: keypoint detection, descriptor extraction, and then matching descriptors from different images.
With the advent of Deep Learning, this separation has become progressively less pronounced. Rather than explicitly designing a keypoint detector and a descriptor, a neural network can be trained directly to produce easily recognizable points and highly discriminative descriptors.
In general, the process can be expressed as
| (8.7) |
where the image is simultaneously transformed into a set of keypoints
and their corresponding descriptors
.
One of the first highly successful algorithms in this family was SuperPoint (DMR18). Its architecture consists of a convolutional network that generates both a probability map indicating keypoint locations and a descriptor for each point. Training is performed in a semi-supervised manner using a technique called Homographic Adaptation, in which artificially warped images are used to improve the stability of the detected features.
Many algorithms have since been proposed to learn feature detection and description jointly. D2-Net (DRP$^+$19) observes that the internal activation maps of a convolutional network naturally contain both geometric and descriptive information, and therefore extracts keypoints directly from the network activations. R2D2 (RWSH19) introduces the concepts of repeatability and reliability, favoring points that are both stable and discriminative. Finally, DISK (TFT20) treats the problem as an optimization of correspondences, seeking to learn directly which points produce the most correct matches.
Matching algorithms have evolved in parallel. In traditional systems, matching is performed using a metric defined on the descriptor, such as Euclidean, Manhattan, or Hamming distance, followed by geometric verification and uniqueness tests. More recent techniques instead seek to learn the matching process directly.
A notable example is SuperGlue (SDMR20), which treats keypoints and their descriptors as graph nodes and uses attention mechanisms to iteratively update the descriptors and produce the final correspondences. Many operations traditionally implemented through cross-checking, uniqueness thresholds, and geometric filtering are thus learned by the network.
LightGlue (LSP23) is a subsequent refinement of this idea. Using adaptive mechanisms, the matching process can be stopped early when the confidence level is sufficiently high, significantly reducing computational cost.
The natural next step in this evolution is LoFTR (Local Feature Transformer) (SSW$^+$21). In this approach, the very concept of a keypoint is called into question: rather than first detecting a set of salient points, the system builds dense image representations and directly searches for correspondences between two observations using transformer-based architectures. This approach is particularly effective in low-texture regions, where traditional corner or blob detectors tend to fail.
This evolution shows how the historical distinction between detection, description, and matching is gradually disappearing. Modern systems increasingly treat the entire problem of image correspondence as a single learning process, in which keypoints, descriptors, and matches emerge jointly during training.
Paolo medici