The Epipolar Plane

The preceding chapters repeatedly noted that a single image cannot provide the world coordinates of the points comprising the image without additional information.

Figure 10.1: Epipolar geometry between two cameras: $\mathbf {t}_1$ and $\mathbf {t_2}$ are the pin-holes, $\mathbf {e}_1$ and $\mathbf {e}_2$ are the epipoles, and the world point $\mathbf {x}$ is projected onto the two image points $\mathbf {p}_1$ and $\mathbf {p}_2$, respectively. All the points involved belong to the same plane.
Image fig_epipolar

Given the pin-hole camera equation (9.26), the only information that a generic image point $\mathbf {p}$ can provide is a relationship among the (infinite) world coordinates $\mathbf {x}$ subtended by the image point, that is, the locus of world coordinates whose projection would produce exactly that particular image point. This relationship is the equation of a line passing through the pin-hole $\mathbf{t}$ and the sensor point corresponding to image point $\mathbf {p}$. Writing equation (9.26) again, it is easy to see the relationship between the parameters of camera i, the image point $\mathbf {p}_i$, and the line representing all possible world points $\mathbf {x}$ subtended by $\mathbf {p}_i$:

\begin{displaymath}
\mathbf{x} = \lambda (\mathbf{K}_{i}\mathbf{R}_{i})^{-1} \ma...
...t}_{i} = \lambda \mathbf{v}_{i}(\mathbf{p}_i) + \mathbf{t}_{i}
\end{displaymath} (10.9)

where $\mathbf{v}_i$ has the same meaning as in equation (9.27), namely, the direction vector from the pin-hole to the sensor point. As follows from the preceding relationship, a single image measurement determines only the optical ray associated with the observed pixel. The three-dimensional point $\mathbf {x}$ is therefore known only up to the depth parameter $\lambda$ along that ray.

In stereo vision, we have two sensors and must therefore define two reference frames with respective parameters $\mathbf{K}_1\mathbf{R}_1$ and $\mathbf{K}_2\mathbf{R}_2$ and pin-hole positions $\mathbf{t_1}$ and $\mathbf {t_2}$, all expressed in world coordinates.

The line (10.9), consisting of the world points associated with image point $\mathbf {p}_1$ seen in the first reference frame, can be projected into the view of the second camera:

\begin{displaymath}
\begin{array}{rl}
\mathbf{p}_2 & = \lambda \mathbf{K}_2 \mat...
...\mathbf{K}^{-1}_1 \mathbf{p}_1 + \mathbf{e}_2 \\
\end{array}
\end{displaymath} (10.10)

where one term varies with the point under consideration and the value $\lambda$, while vector $\mathbf {e}_2$ is always constant and does not depend on the point under consideration.

This constant point is the epipole. The epipole is the intersection point of all epipolar lines and represents the projection of one camera's pin-hole into the other camera's image, that is, the “vanishing point” of the epipolar lines.

For two cameras, the projections of the pin-hole coordinates $\mathbf {t}_1$ and $\mathbf{t}_2$ onto the opposite image are

\begin{displaymath}
\begin{array}{l}
\mathbf{e}_1 = \mathbf{P}_1 \mathbf{t}_2 = ...
...K}_2 \mathbf{R}_2 (\mathbf{t}_1 - \mathbf{t}_2) \\
\end{array}\end{displaymath} (10.11)

where $\mathbf{P}_1$ and $\mathbf{P}_2$ are the projection matrices. Points $\mathbf {e}_1$ and $\mathbf {e}_2$ are the epipoles. If the definitions of the relative poses expressed in (10.4) are substituted into equation (10.11), the image coordinates of the epipoles, understood as the projection of one camera's pin-hole into the other image, are
\begin{displaymath}
\begin{array}{l}
\mathbf{e}_1 = \mathbf{K}_1 \mathbf{R}^{\t...
...thbf{t} \\
\mathbf{e}_2 = \mathbf{K}_2 \mathbf{t}
\end{array}\end{displaymath} (10.12)

functions solely of the relative pose between the two cameras.

By construction, matrix $\mathbf{R}$ converts camera 1 coordinates into camera 2 coordinates, and $\mathbf{t}$ represents the position of camera 1's pin-hole expressed in the reference frame of camera 2.

The lines generated by points in the first image all pass through the same point formed by projecting pin-hole $\mathbf {t}_1$ onto the second image: in fact, the world-coordinate point and the two epipoles define a plane (the epipolar plane) containing the possible solutions—the points in camera coordinates—to the three-dimensional reconstruction problem (Figure 10.1).

Epipolar geometry is the geometry relating two images acquired from different viewpoints. However, the relationships between the images do not depend on the observed scene, but only on the intrinsic parameters of the cameras and their relative poses.

For each observed point, the epipolar plane is the plane defined by the world-coordinate point and the two optical centers. The epipolar line is the intersection between the epipolar plane and the image plane in the second image. In fact, the epipolar plane intersects the two image planes along their respective epipolar lines and constrains the positions of corresponding points in the two images.

The following sections discuss both how to derive the line along which a point belonging to one image must lie in another image and how to obtain the corresponding three-dimensional point from two (or more) corresponding points.

Paolo medici
2026-10-06