Corresponding Points in Alternative Image Spaces

The three-dimensional reconstruction techniques presented so far have been introduced primarily under the pin-hole camera model. This assumption, however, is not strictly necessary. Triangulating a point requires only that each image point be associated uniquely with the corresponding optical ray in three-dimensional space.

In the general case, a camera can be described by a projection function

\begin{displaymath}
\mathbf{p} = \Pi(\mathbf{m})
\end{displaymath} (10.93)

where $\mathbf{m}$ is a point expressed in camera coordinates and $\mathbf {p}$ is its projection onto the image plane. If the inverse function
\begin{displaymath}
\mathbf{v} = \Pi^{-1}(\mathbf{p})
\end{displaymath} (10.94)

makes it possible to recover the direction of the optical ray associated with the observed pixel, then all the geometric relationships introduced in the preceding sections remain valid.

In particular, given two corresponding points observed by two cameras, the corresponding pixels are associated with two optical rays:

\begin{displaymath}
\begin{array}{l}
\mathbf{x} = \mathbf{t}_1 + \lambda_1 \mat...
... \mathbf{x} = \mathbf{t}_2 + \lambda_2 \mathbf{v}_2
\end{array}\end{displaymath} (10.95)

where $\mathbf {t}_1$ and $\mathbf{t}_2$ represent the optical centers of the two cameras, while $\mathbf{v}_1$ and $\mathbf{v}_2$ are the directions associated with the respective image points. Triangulating the three-dimensional point therefore reduces to the same problem discussed in Section 10.3.1, regardless of the camera model used.

This observation makes it possible to extend three-dimensional reconstruction to perspective, Fish-Eye, and omnidirectional cameras, and more generally to any sensor for which the relationship between an observed pixel and the corresponding optical ray is known.

In many practical cases, it is convenient to transform the acquired image into an alternative space using a suitable Warp-Table. The purpose of this transformation is not only to correct the distortion introduced by the optics but also to simplify the search for corresponding points.

To preserve the concept of disparity, and thus continue to apply dense stereo algorithms developed for rectified images, the transformation must keep correspondences aligned on the same image rows. In practical terms, the horizontal coordinate must depend only on the horizontal viewing angle, while the vertical coordinate must depend only on the vertical viewing angle.

A particularly widespread parameterization is the angular (or polar) parameterization, defined by

\begin{displaymath}
\begin{array}{l}
x' = \operatorname{atan} \frac{x}{\sqrt{y^2 + z^2}} \\
y' = \operatorname{atan} \frac{y}{z}
\end{array}\end{displaymath} (10.96)

which transforms a direction vector expressed in camera coordinates $(x,y,z)$ into the alternative image coordinates $(x',y')$.

The inverse transformation is

\begin{displaymath}
\begin{array}{l}
x = \sin(x') \\
y = \cos(x') \sin(y') \\
z = \cos(x') \cos(y')
\end{array}\end{displaymath} (10.97)

and makes it possible to immediately reconstruct the optical ray associated with each pixel in the new representation.

Unlike the classical perspective model, this parameterization can represent points distributed over an extremely wide field of view, up to a complete hemisphere. For this reason, it is frequently used to remap Fish-Eye cameras and panoramic systems.

Once the inverse function associating each pixel with the corresponding three-dimensional optical ray is known, all the techniques discussed in this chapter—including epipolar constraints, Essential Matrix estimation, triangulation, sparse reconstruction, and dense stereo—can be applied without substantial modification. Only the method used to recover the observed optical-ray direction from the image point changes.

Paolo medici
2026-10-01