Change of Viewpoint

The general equation relating image points between two arbitrary viewpoints can be written as

\begin{displaymath}
\begin{bmatrix}
u_2 \\ v_2 \\ 1
\end{bmatrix} \equiv
\mathb...
...n{bmatrix}
u_1 \\ v_1 \\ 1
\end{bmatrix} + \mathbf{t} \right)
\end{displaymath} (9.40)

where $\mathbf{t} = \mathbf{t}_1 - \mathbf{t}_2$ is the vector joining the two pin-holes and $\mathbf{R}$ is the relative orientation between the two views, as described in Section 1.10. A more detailed treatment is provided in Chapter 10 on stereoscopy.

In general, it is not possible to transform a view generated by one camera into the view generated by another. This is possible only when correctly remapping points belonging to a specific plane, or when the cameras share the same pin-hole.

The second case will be discussed in the next section. In the first case, points can be remapped from one view to another by combining a Perspective Mapping followed by an Inverse Perspective Mapping, under the assumption that the observed scene consists only of a plane (for example, the ground). The image points are projected into world coordinates by camera 1 and then reprojected into image coordinates by a second camera, 2, with different intrinsic and extrinsic parameters. Since a plane is always reprojected, the composition of this transformation is also a homography:

\begin{displaymath}
\mathbf{H} = \mathbf{H}_{2} \cdot \mathbf{H}^{-1}_{1}
\end{displaymath} (9.41)

homographic transformations are in fact combined by simply multiplying matrices. Expanding Equation (9.41) with (9.36) yields:
\begin{displaymath}
\mathbf{H} = \mathbf{K}_2 \cdot {\mathbf{R}_{Z}}_{2} \cdot {\mathbf{R}_{Z}}^{-1}_{1} \cdot \mathbf{K}^{-1}_{1}
\end{displaymath} (9.42)

From a theoretical standpoint, forcing plane $z$ to remain constant matters only if the translation vector changes. If the translation vector is modified between the two views and there are points that do not belong to the specified plane, an incorrect remapping occurs between the views (the homographic transformation is no longer valid). Transformation (9.41) can also be used to detect vertical obstacles in techniques such as Ground Plane Stereo and Motion Stereo.

This homography matrix can be generalized by knowing the transformation elements of the two views $(\mathbf{R}, \mathbf{t})$ and the plane equation $(\mathbf{n}, d)$, where $\hat{\mathbf{n}}$ is the normal to the plane and $d$ is the distance between the first camera and the plane. In this case, a point $\mathbf{x}_1$ in the first view belonging to the plane satisfies the equation

\begin{displaymath}
\hat{\mathbf{n}}^{\top} \mathbf{x}_1 = d
\end{displaymath} (9.43)

and this point is related to the same point, as viewed from the second camera, according to the equation
\begin{displaymath}
\mathbf{x}_2 = \mathbf{R} \mathbf{x}_1 + \mathbf{t}
\end{displaymath} (9.44)

Combining these two equations yields the homographic constraint
\begin{displaymath}
\mathbf{H} = \mathbf{K}_2 \left( \mathbf{R} + \frac{1}{d} \mathbf{t} \hat{\mathbf{n}}^{\top} \right) \mathbf{K}^{-1}_{1}
\end{displaymath} (9.45)

A homography can always be decomposed into $\left[ \mathbf{R}, \frac{1}{d} \mathbf{t}, \hat{\mathbf{n}} \right]$ (there are 4 possible decompositions, and the one satisfying the input points must be selected).

Paolo medici
2026-10-01