3D–2D Odometry

The 3D–3D approach is general and can be used with three-dimensional points obtained from any sensor. In stereo vision, however, the three-dimensional points are not direct measurements, but are themselves obtained by triangulating observations from the images.

The noise model is normally formulated for image observations, for which a zero-mean Gaussian error can be assumed as a first approximation. Triangulation, however, transforms this error nonlinearly into the three-dimensional coordinates, making it less straightforward to formulate a maximum-likelihood model operating directly on the reconstructed points.

An alternative approach, referred to as 3D-to-2D, treats the previously reconstructed three-dimensional points as known and estimates the current pose by directly minimizing the reprojection error in the images:

\begin{displaymath}
\sum_i
\left\Vert
\mathbf{p}_i-\hat{\mathbf{p}}_i
\right\Vert^2,
\end{displaymath} (10.103)

where $\hat {\mathbf {p}}_i$ is the projection, after the transformation induced by the motion, of the three-dimensional point $\mathbf{x}_i$ obtained from the preceding frames.

This problem is the same as the Perspective-n-Point (PnP) problem introduced in section 9.5.5. In this particular context, the three-dimensional points $\mathbf{x}_i$ are obtained by triangulating observations acquired in preceding frames, while the pose being estimated represents the relative motion of the camera between instants $t$ and $t'$.

Visual odometry can therefore be interpreted as a sequence of PnP problems, in which an existing three-dimensional map is used to estimate the pose of the current sensor by minimizing the reprojection error.

Paolo medici
2026-10-06