3D–2D Odometry

The 3D–3D approach is general and can be used with three-dimensional points obtained from any sensor. In stereo vision, however, three-dimensional points are not direct measurements; they are themselves obtained by triangulating image observations.

The noise model is normally formulated for image observations, for which a zero-mean Gaussian error can be assumed as a first approximation. Triangulation, on the other hand, transforms this error nonlinearly into the three-dimensional coordinates, making it less straightforward to formulate a maximum-likelihood model that operates directly on the reconstructed points.

An alternative approach is the one known as 3D-to-2D, in which previously reconstructed three-dimensional points are treated as known and the current pose is estimated by directly minimizing the reprojection error in the images:

\begin{displaymath}
\sum_i
\left\Vert
\mathbf{p}_i-\hat{\mathbf{p}}_i
\right\Vert^2,
\end{displaymath} (10.100)

where $\hat {\mathbf {p}}_i$ is the projection, after the motion-induced transformation, of the three-dimensional point $\mathbf{x}_i$ obtained from the preceding frames.

This problem coincides with the Perspective-n-Point (PnP) problem introduced in Section 9.5.5. In this particular context, the three-dimensional points $\mathbf{x}_i$ are obtained by triangulating observations from preceding frames, while the pose being estimated represents the relative motion of the camera between instants $t$ and $t'$.

Visual odometry can therefore be interpreted as a sequence of PnP problems, in which an already known three-dimensional map is used to estimate the pose of the current sensor by minimizing the reprojection error.

Paolo medici
2026-10-01