Descriptor Comparison and Matching

To conclude this chapter, we briefly discuss descriptor comparison.

Let $I_1$ and $I_2$ be two images to be analyzed, and let $\mathbf {p}_1$ and $\mathbf {p}_2$ be two points, likely keypoints, detected in the first and second images, respectively. To determine whether these two image points represent the same physical point—which is typically observed from different viewpoints and therefore affected by affine transformations (translations, scale changes, rotations), homographies, and possibly changes in illumination—we need to define some form of metric $d(\mathbf{p}_1,\mathbf{p}_2)$ for comparison. A particular metric can be defined for each descriptor. The most widely used metrics are the L1 (Manhattan, SAD) and L2 (Euclidean, SSD) distances.

Since more than one point will be extracted from each image, the points must be scanned, and each point in the first image is matched only to the point in the second image with the smallest distance under the selected metric:

\begin{displaymath}
\hat{\mathbf{p}_{2} } = \argmin_i d(\mathbf{p}_1, \mathbf{p}_{2,i} )
\end{displaymath} (8.1)

To reduce the number of incorrect matches, a match is usually accepted only if the distance is below a given threshold and the ratio between the best and second-best matches is below a second, uniqueness threshold.

Finally, after finding $\mathbf {p}_2$, the best match for point $\mathbf {p}_1$ in the second image, we can check whether $\mathbf {p}_2$ has a better match in the first image.

Paolo medici
2026-10-06