Triangulation

Figure 10.2: Triangulation example. Given the camera calibration, the world point $\mathbf {x}$ can be recovered by observing its projection in at least two images ( $\mathbf {p_1}, \mathbf {p_2}, \ldots $). However, because of noise, the resulting lines do not pass through point $\mathbf {x}$ and may not intersect one another. The maximum-likelihood solution requires minimizing the sum of the squared errors between the observed point $\mathbf {p}_i$ and the predicted point $\hat {\mathbf {p}}_i$.
Image fig_triangulate

From Figure 10.2, it is easy to see that the solution to the triangulation problem is the intersection point of the epipolar lines generated by the two images. This problem can readily be extended to the case of $n$ cameras whose relative pose is known. If the relative pose is unknown, it can be obtained, up to a scale factor, directly from the images using techniques such as the Essential matrix (Section 10.4).

Because of inaccuracies in identifying corresponding points (calibration errors could also be considered separately), the lines formed by the optical rays are generally skew. In this case, it is necessary to derive the closest solution under some cost function: the least-squares solution is always possible with $n \geq 2$, using techniques such as Forward Intersections or the Direct Linear Transform (DLT).

Every optical ray subtended by image pixel $(u_i,v_i)$, with $i=1, \ldots, n$ denoting view i, must satisfy equation (10.9). The intersection point (Forward Intersections) of all these rays is the solution of a potentially overdetermined linear system with $3+n$ unknowns in $3n$ equations:

\begin{displaymath}
\left\{
\begin{array}{rl}
\mathbf{x}&= \lambda_1 \mathbf{v}_...
...x}&= \lambda_n \mathbf{v}_n + \mathbf{t}_n
\end{array}\right.
\end{displaymath} (10.13)

where $\mathbf{v}_i = \mathbf{R}^{-1}_i \mathbf{K}^{-1}_i \begin{pmatrix}u_i &v_i&1 \end{pmatrix}^{\top}$ denotes the direction vector of the optical ray in world coordinates. The unknowns are the world point to be estimated, $\mathbf {x}$, and the depth parameters $\lambda_i$ along the respective optical rays.

The closed-form solution, limited to the case of only two lines, is given in Section 1.6.8. This technique can be applied to the case of one camera aligned with the axes and a second camera positioned relative to the first according to relationship (10.5).

Using the properties of the cross product, the same expression can be obtained using perspective projection matrices and image points expressed as homogeneous coordinates:

\begin{displaymath}
\left\{
\begin{array}{l}
\left[ \mathbf{p}_{1} \right]_{\tim...
...right]_{\times} \mathbf{P}_n \mathbf{x} = 0
\end{array}\right.
\end{displaymath} (10.14)

where $[\cdot]_{\times}$ is the cross product written in matrix form. Each of these constraints provides three equations, but only two are linearly independent. All these constraints can finally be rearranged into a homogeneous system of the form
\begin{displaymath}
\mathbf{A} \mathbf{x} = 0
\end{displaymath} (10.15)

where $\mathbf{A}$ is a matrix $2n \times 4$ with $n$ denoting the number of views in which point $\mathbf {x}$ is observed. The solution of homogeneous system (10.15) can be obtained through singular value decomposition. This approach is called the Direct Linear Transform (DLT), by analogy with the calibration technique.

Minimization in world coordinates, however, is not optimal in terms of noise minimization. In the absence of additional information about the structure of the observed scene, the optimal estimate (Maximum Likelihood Estimation) is always the one that minimizes the error in image coordinates (reprojection), but it requires greater computational cost and the use of nonlinear techniques, since the cost function to be minimized is

\begin{displaymath}
\argmin_\mathbf{x} \sum_{i=1}^{n} \Vert \mathbf{p}_i - \hat{\mathbf{p}}_i \Vert^2
\end{displaymath} (10.16)

where $\hat{\mathbf{p}}_i \equiv \mathbf{P}_i \mathbf{x}$ and $\mathbf{P}_i$ is the projection matrix of the i-th image (Figure 10.2).

This is a non-linear, non-convex problem: potentially several local minima are present, and the linear solution must be used as the starting point for the minimization.

An alternative approach consists of estimating the relative pose between two cameras using the Essential matrix. Once this pose is known, the observed three-dimensional points can be reconstructed by triangulation, as shown in Section 10.4.4.

Paolo medici
2026-10-06