Three-Dimensional Reconstruction and Homography

Equation (10.23) can readily be expressed in homogeneous form. The matrix that directly reconstructs the coordinates of a three-dimensional point expressed in the camera reference frame from disparity-image coordinates is

\begin{displaymath}
\begin{bmatrix}
\tilde{x} \\ \tilde{y} \\ \tilde{z} \\ 1
\...
...quiv \mathbf{Q} \begin{bmatrix}
u \\ v \\ d \\ 1
\end{bmatrix}\end{displaymath} (10.24)

whereas its inverse
\begin{displaymath}
\begin{bmatrix}
u \\ v \\ d \\ 1
\end{bmatrix} =
\begin{b...
...matrix}
\tilde{x} \\ \tilde{y} \\ \tilde{z} \\ 1
\end{bmatrix}\end{displaymath} (10.25)

is the matrix that projects a point from camera coordinates to disparity-image coordinates (these matrices are known only up to a multiplicative factor and can therefore be expressed in different forms). The three-dimensional reconstruction of the disparity-image point in the world reference frame, equation (10.19), is equivalent. The matrix $\mathbf{Q}$ is called the reprojection matrix (FK08).

In real conditions, since the camera is rotated and translated with respect to the ideal configuration, it is sufficient to multiply matrix $\mathbf{Q}$ by matrix $4 \times 4$, representing the transformation from camera to world coordinates, to obtain a new matrix that converts disparity coordinates to world coordinates and vice versa.

This formalism makes it possible to transform disparity points acquired by pairs of cameras positioned at different viewpoints (for example, a stereo pair moving over time or two stereo pairs rigidly connected to each other). In this case, the relationship linking disparity points acquired from the two viewpoints is also represented by a matrix $4 \times 4$:

\begin{displaymath}
\mathbf{H}_{2,1} = \mathbf{Q}_1^{-1} \begin{bmatrix}
\math...
...{1}{}{\mathbf{t}}_{2,1} \\
0 & 1
\end{bmatrix} \mathbf{Q}_2
\end{displaymath} (10.26)

which transforms $(u_2,v_2,d_2)$ into $(u_1,v_1,d_1)$ (it is a four-dimensional homographic transformation, entirely analogous to those in three dimensions considered above). Note that the pose $\left( \mathbf{R}, \mathbf{t} \right)$ was used with the syntax of equation (1.100) to express the point from reference frame 2 in reference frame 1. Since all the points involved are expressed in camera coordinates, if transformations between sensors are expressed in world coordinates, as is normally the case, the change of reference frame must also be included. This class of transformations is generally referred to as 3D Homographies.

Paolo medici
2026-10-06