Subsections

Aligned Cameras

For cameras perfectly aligned with the axes and having identical intrinsic parameters (the same focal length and the same principal point), the equations for three-dimensional reconstruction simplify considerably.

Under these conditions, the perspective projection equations reduce to

\begin{displaymath}
\begin{array}{l}
u_i = - k_{u} \dfrac{ y - y_i }{ x - x_i ...
...= - k_{v} \dfrac{ z - z_i }{ x - x_i } + v_{0} \\
\end{array}\end{displaymath} (10.17)

where $(x,y,z)$ is a point in “world” coordinates (see the following section) and $(u_i,v_i)$ are the coordinates of the point projected onto the i-th image. Point $(u_0,v_0)$ is the principal point, which must be the same for all cameras involved, and the cameras are assumed to be perfectly aligned with the three coordinate axes.

Let us now consider only the stereo case: for simplicity, camera 1 will denote the left camera and camera 2 the right camera. The alignment constraints impose $x_1=x_2=0$, $y_1=b$, $y_2=0$, and $z_1=z_2=0$, having placed the right camera at the origin of the reference frame without loss of generality. Quantity $b = y_1 - y_2$ is defined as the baseline.

The difference $d = u_1 - u_2$ between the horizontal coordinates of the projections of the same point viewed in the two images of the stereo pair is called the disparity. This value is obtained by substituting the alignment constraints into equation (10.17), yielding

\begin{displaymath}
u_1 - u_2 = d = k_u \frac{ b }{ x }
\end{displaymath} (10.18)

By inverting this simple relationship and substituting it into equation (10.17), it is possible to obtain the world-coordinate point $(x,y,z)$ corresponding to a point $(u_2,v_2)$ in the right camera with disparity $d$:

\begin{displaymath}
\begin{array}{l}
x = k_u \dfrac{b}{d} \\
y = - (u_2 - u_0...
...
z = - (v - v_0) \dfrac{k_u}{k_v} \dfrac{b}{d} \\
\end{array}\end{displaymath} (10.19)

Clearly, $d \geq 0$ must hold for world points located in front of the stereo pair.

As can be seen, each component is determined by the multiplicative factor $b$ of the baseline, which is the actual scale factor of the reconstruction, and by the inverse of the disparity $1/d$.

Triangulation in World Coordinates

The coordinates $(x,y,z)$ thus obtained are sensor coordinates, referred to a particular stereo configuration in which the orientation and position are aligned with and coincident with the system axes. To move from sensor coordinates to the general case of world coordinates, with arbitrarily oriented cameras, a transformation from sensor to world coordinates must be applied, namely the rotation matrix $\prescript{w}{}{\mathbf{R}}_{b}$ and the translation $(x_i,y_i,z_i)^{\top}$ of the pin-hole, so that

\begin{displaymath}
\begin{bmatrix}
x \\ y \\ z
\end{bmatrix} = \prescript{w...
...nd{bmatrix} + \begin{bmatrix}
x_i \\ y_i \\ z_i
\end{bmatrix}\end{displaymath} (10.20)

By combining equation (10.19) with equation (10.20), it is possible to define a matrix $\mathbf{M}$ such that the conversion between the image-disparity point $(u_i,v,d)$ and the world coordinate $(x,y,z)$ can be written in a very compact form as

\begin{displaymath}
\begin{bmatrix}
x \\ y \\ z
\end{bmatrix} = \frac{1}{d} ...
...nd{bmatrix} + \begin{bmatrix}
x_i \\ y_i \\ z_i
\end{bmatrix}\end{displaymath} (10.21)

where i can denote either the left or the right camera.

Paolo medici
2026-10-01