Subsections


Implicit Calibration

The basic idea of the Direct Linear Transformation proposed by Abdel-Aziz and Karara (AAK71) makes it possible to directly compute the coefficients of matrices (9.56), (9.59), or of matrix (9.25), completely disregarding the parameters and structure of the perspective-transformation model. That paper also presents an approach for solving overdetermined problems using the pseudoinverse technique.

Given system (9.25), it is necessary to recover the 12 parameters of the projective matrix $\mathbf{P}$ to obtain an implicit calibration of the system, that is, one in which the internal parameters (from 9 to 11, depending on the model) that generated the elements of the matrix are unknown. This representation of the pin-hole camera is clearly ideal (with no nonlinearities in the model).

The perspective function written in implicit form is

\begin{displaymath}
\begin{pmatrix}
u_i \\
v_i \\
1
\end{pmatrix} = \mat...
...x} \begin{pmatrix}
x_i \\
y_i \\
z_i \\
1
\end{pmatrix}\end{displaymath} (9.52)

where the elements $p_{0} \ldots p_{11}$ are written in row-major order. It is possible to rearrange system (9.52) to obtain 2 pairs of linear constraints for each point whose image and world coordinates are known:
\begin{displaymath}
\left[ \begin{array}{cccccccccccc}
x_i & y_i & z_i & 1 & 0...
...\begin{pmatrix}
p_0 \\
\vdots \\
p_{11}
\end{pmatrix} = 0
\end{displaymath} (9.53)

This technique is called DLT (direct linear transformation). Since each point provides 2 constraints, at least 6 linearly independent points are required to obtain these 12 parameters; that is, the points must not all lie on the same plane, let alone on the same line.

Since this is a homogeneous system, its solution is the null space of $\mathbb{R}^{12}$, the kernel of the matrix of known terms. For this reason, matrix $\mathbf{P}$ is defined only up to a multiplicative factor, and consequently has only 11 free parameters (there are even fewer when considering that a modern camera has only 3–4 intrinsic parameters and 6 extrinsic parameters). Since the system has been rearranged, noise propagation on the points is no longer linear, and this solution does not satisfy the maximum-likelihood criterion. The matrix $\mathbf{P}$ obtained through this procedure, although it conceals the sensor's internal structure, makes it possible to project a point from world coordinates to image coordinates and, from a point in image coordinates, to recover the line subtended by that point in the world.

The result is generally unstable when only 6 points are used. Therefore, the estimate is normally obtained by processing more than the minimum number of points and using techniques such as the pseudoinverse to determine a solution that minimizes measurement errors.

Generalization to homogeneous coordinates

Equation 9.52 can be generalized to the case of an “image” point in homogeneous coordinates $(u_i, v_i, w_i)$:
\begin{displaymath}
\begin{pmatrix}
u_i \\
v_i \\
w_i
\end{pmatrix} \equ...
...P} \begin{pmatrix}
x_i \\
y_i \\
z_i \\
1
\end{pmatrix}\end{displaymath} (9.54)

The problem is the same as the one considered previously: the homogeneous solution exists, and the homogeneous solution equation (9.53) generalizes to

\begin{displaymath}
{\tiny\left[ \begin{array}{cccccccccccc}
w_i x_i & w_i y_i...
...begin{pmatrix}
p_0 \\
\vdots \\
p_{11}
\end{pmatrix} = 0}
\end{displaymath} (9.55)

for every $i$.

This formulation is useful when the projective model does not follow the pin-hole model, but it is still possible to recover the “camera” coordinates of the optical rays subtended by the pixel, and therefore available in homogeneous form.

DLT computation of the homography

The number of elements in matrix $\mathbf{P}$ can usually be reduced by imposing the constraint that all points involved in the calibration process belong to a particular plane (for example, the ground plane). This means imposing the condition $z_{i}=0$ $\forall i$, which implies removing one column (the one corresponding to axis $z$) from the matrix, which is reduced to size $3 \times 3$, becomes invertible, and can be defined as a homography (see Section 1.11).

We therefore define matrix $\mathbf{H} = \mathbf{P}_Z$ (cf. (9.36)) as

\begin{displaymath}
\lambda
\begin{pmatrix}
u_{i} \\
v_{i} \\
1
\end{pmatrix} = \mathbf{H} \begin{pmatrix}
x_{i} \\
y_{i} \\
1
\end{pmatrix}\end{displaymath} (9.56)

As seen in Section 9.3, this matrix is very useful because, among other things, it makes it possible to remove perspective from the image, synthesizing a fronto-parallel view of the plane through a transformation known as orthogonal rectification, bird's-eye view, or inverse perspective mapping. This transformation can therefore be used both to remove perspective (perspective mapping or inverse perspective mapping), to reproject a plane between two images (ground-plane stereo), and to generate an image with different parameters (rectification, panoramic images) using a virtual plane.

As in the previous case, it is possible to transform the nonlinear relation (9.56) to obtain linear constraints:

\begin{displaymath}
\left[
\begin{array}{ccccccccc}
x_i & y_i & 1 & 0 & 0 & 0 &...
...ight] \begin{pmatrix}
h_0 \\
\vdots \\
h_8
\end{pmatrix} = 0
\end{displaymath} (9.57)

Since this matrix is also defined only up to a multiplicative factor, it has only 8 degrees of freedom, and one additional constraint can therefore be imposed.

If a sufficiently modern linear-system solver is available, the additional constraint $\vert\mathbf{H}\vert=1$ is automatically satisfied while computing the kernel of the matrix of known terms (QR factorization or SVD decomposition).

Another, simpler and more intuitive method consists in imposing the additional constraint $h_{8}=1$. In this way, instead of solving a homogeneous system, a conventional linear problem can be solved. System (9.56) can also be rearranged in this case to obtain linear constraints in the form:

\begin{displaymath}
\left[
\begin{array}{cccccccc}
x_{i} & y_{i} & 1 & 0 & 0 & 0...
...7}
\end{pmatrix}=
\begin{pmatrix}
u_{i} \\
v_{i}
\end{pmatrix}\end{displaymath} (9.58)

This is a (nonhomogeneous) system of two equations in 8 unknowns $h_0 \ldots h_7 $, and each point whose position on a plane in the world and position in the image are both known provides 2 constraints.

However, imposing $h_{8}=1$ means that point $(0,0)$ cannot be an image singularity (e.g., the horizon line), and in general it is not an optimal choice in terms of solution accuracy, as discussed previously.

It is important to note that the solution depends strongly on the chosen normalization. The choice $\vert H\vert=c$ can be called standard least-squares.

In both cases, at least 4 points are required to obtain a homography $\mathbf{H}$, and each additional point makes it possible to obtain a solution with lower error. When overdetermined, these systems can be solved using the pseudoinverse method 1.1.

Matrix $\mathbf{H}$ is defined by 4 intrinsic parameters and 6 extrinsic parameters. Separating the intrinsic parameters from the extrinsic parameters suggests extracting them independently in order to make the calibration more robust. After all, the intrinsic parameters can be determined offline with a certain degree of accuracy and remain valid for all possible camera positions (see 9.5.4 below).

We define matrix $\mathbf{R}_{Z}$ (cf. (9.37)) as

\begin{displaymath}
\lambda
\begin{pmatrix}
\tilde{u_{i}} \\
\tilde{v_{i}} \\
...
...hbf{R}_{Z}
\begin{pmatrix}
x_{i} \\
y_{i} \\
1
\end{pmatrix}\end{displaymath} (9.59)

where $(\tilde{u_{i}},\tilde{v_{i}})$ denotes the so-called normalized image coordinates (homogeneous coordinates of point $(\tilde{x}_i,\tilde{y}_i,\tilde{z}_i)^{\top}$ in camera coordinates).

Matrix $\mathbf{H}$ is defined up to a scale factor, whereas $\mathbf{R}_{Z}$ makes it possible to define the scale because it still contains two orthonormal columns. Knowing the two columns of the rotation matrix makes it possible to recover the third; therefore, this calibration is also valid for points outside plane $z=0$.

As before, a nonlinear system of 3 homogeneous equations, when suitably rearranged, provides two linear constraints:

\begin{displaymath}
\begin{array}{l}
\mathbf{A}\mathbf{x}=0 \\
\mathbf{A} = \be...
...}, r_{6}, r_{7}, p_{x}, p_{y}, p_{z} \right)^{\top}
\end{array}\end{displaymath} (9.60)

(Abdel-Aziz and Karara (AAK71)). It is therefore possible to construct a system of $2 \times N$ equations for all $N$ control points in order to determine the 9 unknowns. The matrix is defined up to a multiplicative factor, but in this case the internal structure of matrix $\mathbf{R}_{Z}$ can help recover the extrinsic parameters (cf. Section 9.5.3). In fact, the two columns of the matrix must be orthonormal:
\begin{displaymath}
\begin{array}{rl}
r^{2}_{0} + r^{2}_{3} + r^{2}_{6} & = 1 ...
...\
r_{0}r_{1} + r_{3}r_{4} + r_{6}r_{7} & = 0 \\
\end{array}\end{displaymath} (9.61)

These additional nonlinear constraints result from the fact that this matrix is explicitly defined by only 6 parameters (3 rotations and the translation).

Geometric representation

Equations (9.53) and (9.57) can also be derived from purely geometric considerations, since the image and camera vectors must be parallel (factor $lamba_i$ is purely multiplicative and can affect the vector only through an affine transformation):

\begin{displaymath}
\mathbf{p} \times \mathbf{P} \mathbf{x} = \mathbf{0}
\qquad
\mathbf{m}' \times \mathbf{H} \mathbf{m} = \mathbf{0}
\end{displaymath} (9.62)

This compact formulation is normally denoted as DLT (HZ04) and applies to all linear transformations known only up to a multiplicative factor, converting the problem into a homogeneous one.

Paolo medici
2026-10-01