Perspective Mapping and Inverse Perspective Mapping

Using a homography, it is possible to implement the inverse perspective mapping (or bird eye view) transformation by simply inverting the perspective mapping matrix.

The homography matrix $\mathbf{H} = \mathbf{P}_{Z}$ of the perspective projection of a plane, perspective mapping, for a plane $z$ held constant, where normally $z=0$ since the ground is the most important plane, can be derived very simply because:

\begin{displaymath}
\mathbf{P}_{Z} = \mathbf{K} \cdot \mathbf{R}_{Z}
\end{displaymath} (9.36)

where $\mathbf{R}_{Z}$ is the rigid transformation matrix of a plane, which can be expressed as
\begin{displaymath}
\mathbf{R}_{Z} = \begin{bmatrix}\mathbf{r}_1 & \mathbf{r}_2...
...de{t}_y \\
r_6 & r_7 & r_8 z + \tilde{t}_z \\
\end{bmatrix}\end{displaymath} (9.37)

where vector $\mathbf{\tilde{t}}$ denotes the translation expressed in camera coordinates, as in Equation (9.23).

This matrix is very important and will be discussed extensively in Section 9.5 on calibration.

Transformation (9.36), being a homography, is invertible. When it densely transforms all image points into world points, it is called Inverse Perspective Mapping; when it transforms all world points into image points, it is called Perspective Mapping. In both cases, only plane $z$ is projected correctly.

It is worth noting that even the simplest 9-parameter pin-hole camera model (6 extrinsic and 3 intrinsic parameters) cannot be recovered from the 8 constraint parameters provided by the homography matrix. However, if the intrinsic parameters are known, it is possible to obtain an estimate of the camera's rotation and position (Section 9.5), since Equation 9.36 becomes invertible:

\begin{displaymath}
\mathbf{R}_{Z} = \mathbf{K}^{-1} \mathbf{H}
\end{displaymath} (9.38)

Paolo medici
2026-10-01