Subsections

Aligned Cameras

For cameras perfectly aligned with the axes and having equal intrinsic parameters (the same focal length and the same principal point), the equations for three-dimensional reconstruction simplify considerably.

Under these conditions, the perspective projection equations reduce to

\begin{displaymath}
\begin{array}{l}
u_i = - k_{u} \dfrac{ y - y_i }{ x - x_i ...
...= - k_{v} \dfrac{ z - z_i }{ x - x_i } + v_{0} \\
\end{array}\end{displaymath} (10.17)

where $(x,y,z)$ is a point in “world” coordinates (see the following section) and $(u_i,v_i)$ are the coordinates of the point projected onto the i-th image. Point $(u_0,v_0)$ is the principal point, which must be the same for all cameras involved; by assumption, these cameras must be perfectly aligned with the coordinate axes.

We now restrict ourselves to the stereo case: for simplicity, the left camera will be denoted by subscript 1 and the right camera by 2. The alignment constraints impose $x_1=x_2=0$, $y_1=b$, $y_2=0$, and $z_1=z_2=0$, having placed the right camera at the origin of the reference frame without loss of generality. Quantity $b = y_1 - y_2$ is called the baseline.

The difference $d = u_1 - u_2$ between the horizontal coordinates of the projections of the same point viewed in the two images of the stereo pair is called the disparity. This value is obtained by inserting the alignment constraints into equation (10.17), yielding

\begin{displaymath}
u_1 - u_2 = d = k_u \frac{ b }{ x }
\end{displaymath} (10.18)

Inverting this simple relationship and substituting it into equation (10.17) makes it possible to recover the world-coordinate point $(x,y,z)$ corresponding to a point $(u_2,v_2)$ in the right camera with disparity $d$:

\begin{displaymath}
\begin{array}{l}
x = k_u \dfrac{b}{d} \\
y = - (u_2 - u_0...
...
z = - (v - v_0) \dfrac{k_u}{k_v} \dfrac{b}{d} \\
\end{array}\end{displaymath} (10.19)

forcing $v=v_1=v_2$. Clearly, $d \geq 0$ must hold for world points located in front of the stereo pair.

As can be seen, each coordinate is determined by the multiplicative factor $b$ of the baseline, which is the actual scale factor of the reconstruction, and by the inverse of the disparity $1/d$.

Triangulation in World Coordinates

The coordinates $(x,y,z)$ thus obtained are sensor coordinates, referred to a particular stereo configuration in which orientation and position are aligned with and coincident with the coordinate axes. To move from sensor coordinates to the general case of world coordinates, with arbitrarily oriented cameras, a transformation from sensor to world coordinates must be applied, namely the rotation matrix $\prescript{w}{}{\mathbf{R}}_{b}$ and the translation $(x_i,y_i,z_i)^{\top}$ of the pin-hole coordinates, so that we can write

\begin{displaymath}
\begin{bmatrix}
x \\ y \\ z
\end{bmatrix} = \prescript{w...
...nd{bmatrix} + \begin{bmatrix}
x_i \\ y_i \\ z_i
\end{bmatrix}\end{displaymath} (10.20)

Combining equation (10.19) with equation (10.20), it is possible to define a matrix $\mathbf{M}$ such that the conversion between the disparity-image point $(u_i,v,d)$ and the world coordinate $(x,y,z)$ can be written in the very compact form

\begin{displaymath}
\begin{bmatrix}
x \\ y \\ z
\end{bmatrix} = \frac{1}{d} ...
...nd{bmatrix} + \begin{bmatrix}
x_i \\ y_i \\ z_i
\end{bmatrix}\end{displaymath} (10.21)

where i can denote either the left or the right camera.

Paolo medici
2026-10-06