Matching and Learned Local Features

The methods described in the preceding chapters typically involve three distinct stages: detecting keypoints, extracting the descriptor, and finally matching descriptors from different images.

With the advent of Deep Learning, this separation has gradually become less pronounced. Instead of explicitly designing a keypoint detector and a descriptor, a neural network can be trained directly to produce easily recognizable points and highly discriminative descriptors.

In a general formulation, the process can be expressed as


\begin{displaymath}
(\mathcal{K},\mathcal{D}) = f(I)
\end{displaymath} (8.7)

where image $I$ is simultaneously transformed into the set of keypoints $\mathcal{K}$ and their corresponding descriptors $\mathcal{D}$.

One of the first highly successful algorithms in this family was SuperPoint (DMR18). Its architecture consists of a convolutional network that generates both a probability map indicating the locations of keypoints and a descriptor associated with each point. Training is performed in a semi-supervised manner using a technique called Homographic Adaptation, in which artificially warped images are used to improve the stability of the detected features.

Numerous algorithms have subsequently been proposed to jointly learn feature detection and description. D2-Net (DRP$^+$19) observes that the internal activation maps of a convolutional network naturally contain both geometric and descriptive information, and therefore extracts keypoints directly from the network activations. R2D2 (RWSH19) instead introduces the concepts of repeatability and reliability, favoring points that are both stable and discriminative. Finally, DISK (TFT20) treats the problem as a correspondence-optimization process, seeking to learn directly which points produce the greatest number of correct matches.

In parallel, matching algorithms have also evolved. In traditional systems, matching is performed using a metric defined on the descriptor, such as Euclidean, Manhattan, or Hamming distance, followed by geometric verification and uniqueness tests. More recent techniques instead seek to learn the matching process directly.

A significant example is SuperGlue (SDMR20), which treats keypoints and their descriptors as the nodes of a graph and uses attention mechanisms to iteratively update the descriptions and produce the final correspondences. Many operations traditionally implemented through cross-checking, uniqueness thresholds, and geometric filtering are thus learned by the network.

LightGlue (LSP23) represents a subsequent refinement of this idea. Through adaptive mechanisms, the matching process can be terminated early when the attained confidence level is sufficiently high, significantly reducing the computational cost.

The natural evolution of this line of work is LoFTR (Local Feature Transformer) (SSW$^+$21). In this case, the entire concept of a keypoint is called into question: instead of first identifying a set of salient points, the system constructs dense image representations and directly searches for correspondences between two observations using transformer-based architectures. This approach is particularly effective in regions with little texture, where traditional corner or blob detectors tend to fail.

The evolution described above shows how the historical distinction between detection, description, and matching is progressively disappearing. Indeed, modern systems tend to treat the entire problem of establishing correspondences between images as a single learning process, in which keypoints, descriptors, and matching emerge jointly during training.

Paolo medici
2026-10-01