Essential and Fundamental Matrices

In 1981, Christopher Longuet-Higgins (Lon81) was the first to observe that a generic point expressed in world coordinates, its corresponding points in camera coordinates, and the pin-holes must be coplanar. The geometric derivation of the relationships between the points is omitted; the analytical derivation is presented directly.

It has been stated repeatedly that a point in an image subtends a line in the world, and that the line in the world projected onto another image, acquired by a camera at a different viewpoint, is the epipolar line on which the point corresponding to the point in the first image lies. This equation, which relates points in one image to lines in the other, can be expressed in matrix form.

Following Higgins's reasoning, the intrinsic-parameter matrix will be omitted, and the coordinates used will be normalized camera coordinates.

Without loss of generality, consider a system consisting of two cameras, with the first positioned and oriented with respect to the second by the projection matrix $\mathbf{P}_1=[\mathbf{R}\vert\mathbf{t}]$, while the second is located at the origin of the reference frame and aligned with the axes, with projection matrix $\mathbf{P}_2=[\mathbf{I}\vert\mathbf{0}]$. The same result can be obtained starting from two generic calibrated cameras, arbitrarily oriented and positioned with respect to a third system, through relations $\mathbf{R} = \prescript{2}{}{\mathbf{R}}_{1} = \mathbf{R}^{-1}_{2} \mathbf{R}_{1}$ and $\mathbf{t}$, namely, the position of camera 1 with respect to system 2.

A generic point $\mathbf{x} \in \mathbb{R}^3$ has coordinates $\mathbf{x}_1$ and $\mathbf{x}_2$ in the two different reference frames and is projected onto sensors 1 and 2 at the points with camera coordinates $\mathbf{m}_1$ and $\mathbf{m}_2$, respectively.

We know that these image points subtend a subspace of $\mathbb{R}^3$, with equation, for example, $\lambda \mathbf{m}_2$, passing through the pin-hole of the second sensor (here constrained to be at $\mathbf{0}$), namely

\begin{displaymath}
\mathbf{x}_2 = \lambda_2 \mathbf{m}_2 + \mathbf{0}
\end{displaymath} (10.38)

A generic point $\mathbf{x}_1 = \lambda_1 \mathbf{m}_1$ expressed in the coordinates of sensor 1 and observed by that sensor can be projected into the coordinates of sensor 2 according to

\begin{displaymath}
\mathbf{x}_2 = \prescript{2}{}{ \mathbf{R}_1 } \mathbf{x}_1 + \mathbf{t}
\end{displaymath} (10.39)

The equation of the epipolar line—the line in camera coordinates in the second sensor and the locus of points on which $\mathbf{m}_2$ must lie—associated with point $\mathbf{m}_1$ (observed and therefore expressed in the camera coordinates of the first sensor) is

\begin{displaymath}
\lambda_2 \mathbf{m}_2 = \lambda_1 \prescript{2}{}{ \mathbf{R}_1 } \mathbf{m}_1 + \mathbf{t}
\end{displaymath} (10.40)

The locus of homogeneous points $\mathbf{m}_2$ is obtained by varying parameter $\lambda_1$, and this line in $\mathbb{R}^3$ is also a line in $\mathbb{R}^2$. If two points are actually corresponding points, the system is solvable and the parameters $\lambda_1$ and $\lambda_2$ can be recovered (this is an example of three-dimensional reconstruction by triangulation, as discussed in Section 10.3.1).

There is, however, a relation between the points in the two cameras that eliminates the parameters $\lambda$ and, above all, makes it possible to reverse the reasoning, namely, to recover the relative pose between the two cameras $\left( \mathbf{R}, \mathbf{t} \right)$ from a set of corresponding points.

If both sides of equation (10.40) are multiplied, first vectorially by $\mathbf{t}$ and then scalarly by $\mathbf{m}^{\top}_{2}$, one obtains

\begin{displaymath}
\lambda_2 \mathbf{m}^{\top}_{2} \left( \mathbf{t} \times \ma...
...thbf{m}^{\top}_{2} \left( \mathbf{t} \times \mathbf{t} \right)
\end{displaymath} (10.41)

The properties of the vector product $\mathbf{t} \times \mathbf{t}=\textbf{0}$ and scalar product $\mathbf{m}_2 \cdot \left( \mathbf{t} \times \mathbf{m}_2 \right)=0$ can be applied to this relation.

This step has a physical interpretation: the coplanarity constraints are introduced first (all expressed, for example, in reference frame 2) among points $\mathbf{0}$ (the pinhole of camera 2), $\mathbf{m}_2$, $\prescript{2}{}{ \mathbf{R}_1 } \mathbf{m}_1 + \mathbf{t}$, $\mathbf{x}_2 = \prescript{2}{}{ \mathbf{R}_1 } \mathbf{x}_1 + \mathbf{t}$, and $\mathbf{t}$ (the pinhole of camera 1 in system 2), and combined with the fact that the body is rigid.

Using this formula, the relationships between corresponding points $\mathbf{m}_1$ and $\mathbf{m}_2$, represented as homogeneous camera coordinates, can be expressed in the compact form

\begin{displaymath}
\mathbf{m}^{\top}_{2} \left( \mathbf{t} \times \mathbf{ R } \mathbf{m}_1 \right) = 0
\end{displaymath} (10.42)

Finally, denoting by $[\mathbf{t}]_{\times}$ the skew-symmetric matrix representing the vector product in matrix form (Section 1.8), the various contributions can be collected into the matrix
\begin{displaymath}
\mathbf{E} = [\mathbf{t}]_{\times} \mathbf{R} = \mathbf{R} \left[ \mathbf{R}^{\top} \mathbf{t} \right]_{\times}
\end{displaymath} (10.43)

where the meaning of the matrices must be carefully recalled by comparison with equation (10.39). A linear relation linking the camera points in the two views can thus be defined:
\begin{displaymath}
\mathbf{m}^{\top}_{2} \mathbf{ E } \mathbf{m}_1 = 0
\end{displaymath} (10.44)

Matrix $\mathbf{E}$ is called the Essential Matrix.

Finally, great care must be taken with the indices, because there is no unique convention for denoting points 1 and 2: assuming convention (10.43), what must be remembered is that matrix $\mathbf{E}$ encodes the relative pose of the camera of the points on the right (in our case $\mathbf{m}_1$) of matrix $\mathbf{E}$ with respect to the camera of the points on the left (in our case $\mathbf{m}_2$) of the matrix.

Matrix $\mathbf{E}$, relating homogeneous points, is also homogeneous and is therefore defined up to a multiplicative factor.

The Essential Matrix has the following properties:

defined as the image of the other camera centre)

The Essential Matrix defines relations in camera coordinates and therefore, to use it in practice, the points must be available in this particular reference frame; that is, the intrinsic parameters of the cameras involved must be known.

Equation

\begin{displaymath}
\mathbf{m}^{\top}_{2} \left( \mathbf{ E } \mathbf{m}_1 \right) = 0
\end{displaymath} (10.45)

can also be interpreted as the equation of a plane in space $2$ passing through $\mathbf{0}$, namely, the epipolar plane formed by the two epipoles and the world point, the plane to which point $\mathbf{m}_{2}$ must belong.

It is nevertheless possible to introduce an additional relation between image points while completely disregarding the cameras' intrinsic parameters.

Applying the definition of homogeneous camera coordinates $\mathbf{p} = \mathbf{K} \mathbf{m}$ in relation (10.44) yields

\begin{displaymath}
\mathbf{m}^{\top}_2 \mathbf{E} \mathbf{m}_1 =
\mathbf{p}^{\t...
...1} \mathbf{p}_1 =
\mathbf{p}^{\top}_2 \mathbf{F} \mathbf{p}_1
\end{displaymath} (10.46)

The Fundamental Matrix is defined (Faugeras and Hartley, 1992) as:

\begin{displaymath}
\mathbf{p}^{\top}_{2} \mathbf{ F } \mathbf{p}_1 = 0
\end{displaymath} (10.47)

where $\mathbf {p}_1$ and $\mathbf {p}_2$ are homogeneous coordinates of the corresponding points in the first and second images, respectively.

If two points in the two images of the stereo pair represent the same world point, equation (10.47) must be satisfied.

The Fundamental Matrix narrows the search interval for correspondences between the two images because, by point-line duality, the locus of points in the second image where the point from the first image must be sought can be made explicit from relation (10.47). Indeed, the equation of a line on which points $\mathbf{m_2}$ and $\mathbf{m_1}$ must lie is given by

\begin{displaymath}
\begin{array}{l}
\mathbf{l}_2 = \mathbf{F} \mathbf{m}_1 \\
\mathbf{l}_1 = \mathbf{F}^{\top} \mathbf{m}_2 \\
\end{array}\end{displaymath} (10.48)

where $\mathbf{l}_1$ and $\mathbf{l}_2$ are the parameters of the epipolar line belonging to the first and second image, respectively, written in implicit form.

The relationship between the Fundamental Matrix and the Essential Matrix is, according to equation (10.46),

\begin{displaymath}
\mathbf{E} = \mathbf{K}^{\top}_{2} \mathbf{F} \mathbf{K}_{1}
\end{displaymath} (10.49)

or conversely,
\begin{displaymath}
\mathbf{F} = \mathbf{K}^{-\top}_{2} \mathbf{E} \mathbf{K}^{...
...\top}_{2} [\mathbf{t}]_{\times} \mathbf{R} \mathbf{K}^{-1}_{1}
\end{displaymath} (10.50)

The Essential Matrix encodes the relative poses of the cameras, whereas the Fundamental Matrix conceals both the intrinsic parameters and the relative pose.

The Essential Matrix introduces constraints identical to those of the Fundamental Matrix but, although it was historically introduced before the Fundamental Matrix, it is a special case of it because it expresses the relations in camera coordinates.

$\mathbf{F}$ is a $3 \times 3$ rank-2 matrix, and 7 points are sufficient to determine it, since it has exactly 7 degrees of freedom (a multiplicative factor and the zero determinant reduce the dimensionality of the problem). The relationship between the Fundamental Matrix and its 7 degrees of freedom is nonlinear (it cannot be readily expressed through a simple algebraic representation). With (at least) 8 points, however, a linear estimate of the matrix can be obtained, as described in the following section.

The Fundamental Matrix has the following properties:

Figure 10.3: The Fundamental Matrix identifies the epipolar lines in the right image on which the corresponding points of points in the left image lie.
Image fig_fundamental

The Fundamental and Essential Matrices can be used to narrow the search for corresponding points between two images and/or to filter out possible outliers (for example, in RANSAC). When decomposed, the Essential Matrix makes it possible to recover the relative pose between the two cameras and, as such, to obtain an approximate estimate of the motion undergone by a camera moving through the world (motion stereo) or of the relative pose of two cameras in a stereo pair (self-calibration).

Using the Essential Matrix makes it possible to recover the relative pose between two views. However, the length of the baseline connecting the two pin-holes cannot be determined; only its direction can be recovered. Nevertheless, given the Essential Matrix, it is always possible to perform a three-dimensional reconstruction of the observed scene up to a multiplicative factor: the ratios between distances are known, but not their absolute values. When the same scene is observed from more than two different views, this makes it possible to obtain a consistent three-dimensional reconstruction in which the unknown multiplicative factor remains the same for all views, thereby allowing all individual reconstructions to be fused into a single reconstruction known up to the same scale factor.



Subsections
Paolo medici
2026-10-01