Three-Dimensional Reconstruction and Homography

Equation (10.23) can easily be expressed in homogeneous form. The matrix that directly reconstructs, from image-disparity coordinates, the coordinates of the three-dimensional point expressed in the camera reference frame is

\begin{displaymath}
\begin{bmatrix}
\tilde{x} \\ \tilde{y} \\ \tilde{z} \\ 1
\...
...x} = \mathbf{Q} \begin{bmatrix}
u \\ v \\ d \\ 1
\end{bmatrix}\end{displaymath} (10.24)

whereas its inverse
\begin{displaymath}
\begin{bmatrix}
u \\ v \\ d \\ 1
\end{bmatrix} =
\begin{b...
...matrix}
\tilde{x} \\ \tilde{y} \\ \tilde{z} \\ 1
\end{bmatrix}\end{displaymath} (10.25)

is the matrix that projects a point from camera coordinates into image-disparity coordinates (these matrices are known up to a multiplicative factor and can therefore be expressed in different forms). The three-dimensional reconstruction of the image-disparity point in the world reference frame, equation (10.19), is equivalent. The matrix $\mathbf{Q}$ is called the reprojection matrix (FK08).

In real conditions, since the camera is rotated and translated with respect to the ideal configuration, it is sufficient to multiply matrix $\mathbf{Q}$ by matrix $4 \times 4$, representing the transformation from camera to world coordinates, to obtain a new matrix that converts disparity coordinates into world coordinates and vice versa.

This formalism can be used to transform disparity points acquired by pairs of cameras positioned at different viewpoints (for example, a stereo pair moving over time or two stereo pairs rigidly connected to each other). In this case, the relation between disparity points acquired from the two viewpoints is also represented by a matrix $4 \times 4$:

\begin{displaymath}
\mathbf{H}_{2,1} = \mathbf{Q}_1^{-1} \begin{bmatrix}
\math...
...{1}{}{\mathbf{t}}_{2,1} \\
0 & 1
\end{bmatrix} \mathbf{Q}_2
\end{displaymath} (10.26)

which transforms $(u_2,v_2,d_2)$ into $(u_1,v_1,d_1)$ (it is a four-dimensional homographic transformation, entirely analogous to those in three dimensions discussed above). Note that the pose $\left( \mathbf{R}, \mathbf{t} \right)$ was written using the syntax of equation (1.100), expressing the point from reference frame 2 in reference frame 1. Since all the points involved are expressed in camera coordinates, if the transformations between sensors are expressed in world coordinates, as is normally the case, the change of reference frame must also be included. This class of transformations is commonly referred to as 3D Homographies.

Paolo medici
2026-10-01