In 1981, Christopher Longuet-Higgins (Lon81) was the first to observe that a generic point expressed in world coordinates, its corresponding points in camera coordinates, and the pin-holes must be coplanar. The geometric derivation of the relationships between the points is omitted; the analytical derivation is presented directly.
It has been stated repeatedly that a point in an image subtends a line in the world, and that the line in the world projected onto another image, acquired by a camera at a different viewpoint, is the epipolar line on which the point corresponding to the point in the first image lies. This equation, which relates points in one image to lines in the other, can be expressed in matrix form.
Following Higgins's reasoning, the intrinsic-parameter matrix will be omitted, and the coordinates used will be normalized camera coordinates.
Without loss of generality, consider a system consisting of two cameras, with the first positioned and oriented with respect to the second by the projection matrix
, while the second is located at the origin of the reference frame and aligned with the axes, with projection matrix
. The same result can be obtained starting from two generic calibrated cameras, arbitrarily oriented and positioned with respect to a third system, through relations
and
, namely, the position of camera 1 with respect to system 2.
A generic point
has coordinates
and
in the two different reference frames and is projected onto sensors 1 and 2 at the points with camera coordinates
and
, respectively.
We know that these image points subtend a subspace of , with equation, for example,
, passing through the pin-hole of the second sensor (here constrained to be at
), namely
| (10.38) |
A generic point
expressed in the coordinates of sensor 1 and observed by that sensor can be projected into the coordinates of sensor 2 according to
The equation of the epipolar line—the line in camera coordinates in the second sensor and the locus of points on which must lie—associated with point
(observed and therefore expressed in the camera coordinates of the first sensor) is
There is, however, a relation between the points in the two cameras that eliminates the parameters and, above all, makes it possible to reverse the reasoning, namely, to recover the relative pose between the two cameras
from a set of corresponding points.
If both sides of equation (10.40) are multiplied, first vectorially by and then scalarly by
, one obtains
| (10.41) |
This step has a physical interpretation: the coplanarity constraints are introduced first (all expressed, for example, in reference frame 2) among points (the pinhole of camera 2),
,
,
, and
(the pinhole of camera 1 in system 2), and combined with the fact that the body is rigid.
Using this formula, the relationships between corresponding points and
, represented as homogeneous camera coordinates, can be expressed in the compact form
Finally, great care must be taken with the indices, because there is no unique convention for denoting points 1 and 2: assuming convention (10.43), what must be remembered is that matrix encodes the relative pose of the camera of the points on the right (in our case
) of matrix
with respect to the camera of the points on the left (in our case
) of the matrix.
Matrix , relating homogeneous points, is also homogeneous and is therefore defined up to a multiplicative factor.
The Essential Matrix has the following properties:
defined as the image of the other camera centre)
The Essential Matrix defines relations in camera coordinates and therefore, to use it in practice, the points must be available in this particular reference frame; that is, the intrinsic parameters of the cameras involved must be known.
Equation
| (10.45) |
It is nevertheless possible to introduce an additional relation between image points while completely disregarding the cameras' intrinsic parameters.
Applying the definition of homogeneous camera coordinates
in relation (10.44) yields
The Fundamental Matrix is defined (Faugeras and Hartley, 1992) as:
If two points in the two images of the stereo pair represent the same world point, equation (10.47) must be satisfied.
The Fundamental Matrix narrows the search interval for correspondences between the two images because, by point-line duality, the locus of points in the second image where the point from the first image must be sought can be made explicit from relation (10.47).
Indeed, the equation of a line on which points and
must lie is given by
| (10.48) |
The relationship between the Fundamental Matrix and the Essential Matrix is, according to equation (10.46),
| (10.49) |
| (10.50) |
The Essential Matrix encodes the relative poses of the cameras, whereas the Fundamental Matrix conceals both the intrinsic parameters and the relative pose.
The Essential Matrix introduces constraints identical to those of the Fundamental Matrix but, although it was historically introduced before the Fundamental Matrix, it is a special case of it because it expresses the relations in camera coordinates.
is a
rank-2 matrix, and 7 points are sufficient to determine it, since it has exactly 7 degrees of freedom (a multiplicative factor and the zero determinant reduce the dimensionality of the problem).
The relationship between the Fundamental Matrix and its 7 degrees of freedom is nonlinear (it cannot be readily expressed through a simple algebraic representation).
With (at least) 8 points, however, a linear estimate of the matrix can be obtained, as described in the following section.
The Fundamental Matrix has the following properties:
|
The Fundamental and Essential Matrices can be used to narrow the search for corresponding points between two images and/or to filter out possible outliers (for example, in RANSAC). When decomposed, the Essential Matrix makes it possible to recover the relative pose between the two cameras and, as such, to obtain an approximate estimate of the motion undergone by a camera moving through the world (motion stereo) or of the relative pose of two cameras in a stereo pair (self-calibration).
Using the Essential Matrix makes it possible to recover the relative pose between two views. However, the length of the baseline connecting the two pin-holes cannot be determined; only its direction can be recovered. Nevertheless, given the Essential Matrix, it is always possible to perform a three-dimensional reconstruction of the observed scene up to a multiplicative factor: the ratios between distances are known, but not their absolute values. When the same scene is observed from more than two different views, this makes it possible to obtain a consistent three-dimensional reconstruction in which the unknown multiplicative factor remains the same for all views, thereby allowing all individual reconstructions to be fused into a single reconstruction known up to the same scale factor.