World Coordinates and Camera Coordinates

Figure 9.2: Image Coordinates (Image coordinates)
Figure 9.3: Camera Coordinates (Camera coordinates)
Figure 9.4: Example of “Vehicle” or “World” coordinates: Front-Left-Up or ISO 8855 (World coordinates)
Image fig_imagecoord

Image fig_cameracoord

Image fig_worldcoord

When dealing with practical problems, it is necessary to move from a reference frame attached to the camera, where point $(0,0,0)^{\top}$ coincides with the focus (pin-hole), to a more general reference frame that better suits the user's requirements, where the camera is located at a generic point in the “world” and oriented arbitrarily with respect to it. This applies to any generic sensor, including non-video sensors, by defining relationships that make it possible to transform points from world coordinates to sensor coordinates and vice versa.

At this point, a clarification is necessary concerning the terminology related to reference frames used in this book: the “world” reference frame is defined as the frame that, in each context, is considered absolute and fixed, with respect to which the sensor is positioned. For example, in Figure 9.4, the origin of the “world” frame is associated with a point on the vehicle (the front point, for example). In this case, the “vehicle” (body) and “world” (world) frames are synonymous. This distinction disappears, however, when a vehicle moves with respect to a “world” that can once again be defined as the fixed reference frame. In that case, we have sensor coordinates, the local vehicle/body coordinates, and finally world coordinates. Usually, however, the axis convention distinguishing the sensor, vehicle, and world frames is kept consistent.

If the special role of coordinate $z$ in camera coordinates is due to purely mathematical reasons, namely the use of homogeneous coordinates, which during projection requires the first two components to be divided by the third, this restriction does not apply to “sensor” coordinates. Although not imposed in any way, this book uses the system shown in Figure 9.4 (ISO 8855) as the “sensor”, “body”, and “world” frame, assigning the height of the point above the ground to axis $z$.

Therefore, to derive the final equation of the pin-hole camera, we start from Equation (9.10) and apply the following considerations:

The conversion from “world” coordinates to “camera” coordinates, being a composition of rotations, is itself a rotation described by equation $\mathbf{R} = \prescript{c}{}{\mathbf{R}}_{w} = \boldsymbol\Pi \prescript{w}{}{\mathbf{R}}^{-1}_{b}$.

Let $(x_{i},y_{i},z_{i})^{\top}$ be a point in “world” coordinates and $(\tilde{x}_{i},\tilde{y}_{i},\tilde{z}_{i})^{\top}$ the same point in “camera” coordinates. The relationship between these two points can be written as

\begin{displaymath}
\begin{bmatrix}
\tilde{x}_{i} \\
\tilde{y}_{i} \\
\tilde{z...
..._{i} \\
y_{i} \\
z_{i}
\end{bmatrix} + \tilde{\mathbf{t}}_0
\end{displaymath} (9.22)

where $\mathbf{R}$ is a $3 \times 3$ matrix that converts world coordinates to camera coordinates and accounts for the rotations and the change in the sign of the axes between world and camera coordinates (see Appendix A), while the vector
\begin{displaymath}
\tilde{\mathbf{t}}_{0} = -\mathbf{R} \mathbf{t}_0
\end{displaymath} (9.23)

represents the position of the pin-hole $\mathbf{t}_0$ with respect to the origin of the world frame, expressed in the camera coordinate system.

Recall that rotation matrices are orthonormal matrices: they have determinant 1 and therefore preserve distances and areas, and the inverse of a rotation matrix is its transpose.

Matrix $\mathbf{R}$ and vector $\mathbf{t}_{0}$ can be combined into a $3\times4$ matrix by exploiting homogeneous coordinates. This representation makes it possible to write, in an extremely compact form, the projection of a point expressed in homogeneous world coordinates, $(x_{i},y_{i},z_{i})^{\top}$, onto an image point with homogeneous coordinates $(u_{i},v_{i})^{\top}$:

\begin{displaymath}
\lambda \begin{bmatrix}
u_{i} \\
v_{i} \\
1
\end{bmatri...
...] \begin{bmatrix}
x_{i} \\
y_{i} \\
z_{i} \\
1
\end{bmatrix}\end{displaymath} (9.24)

This equation makes it clear that each image point $(u_{i},v_{i})$ is associated with infinitely many world points $(x_{i},y_{i},z_{i})^{\top}$ lying on a line as parameter $\lambda$ varies.

Suppressing $\lambda$ and collecting the matrices yields the final equation of the pin-hole camera (which does not and must not account for distortion):

\begin{displaymath}
\begin{bmatrix}
u_{i} \\
v_{i} \\
1
\end{bmatrix} = \mathb...
...} \begin{bmatrix}
x_{i} \\
y_{i} \\
z_{i} \\
1
\end{bmatrix}\end{displaymath} (9.25)

where $\mathbf{P} = \mathbf{K} [ \mathbf{R} \vert \mathbf{\tilde{t}}_0 ]$ has been defined as the projective matrix (camera matrix), which will be used later (Str87). Matrix $\mathbf{P}$ is a $3\times4$ matrix and, being rectangular, is not invertible.

It should be noted that by imposing an additional constraint on the points, for example $z_{i}=0$, matrix $\mathbf{P}$ reduces to an invertible $3 \times 3$ matrix, which is exactly the homography matrix (see Section 9.3.1) of the perspective transformation of ground points. Matrix $\mathbf{P}_{z=0}$ is an example of an IPM (Inverse Perspective Mapping) transformation for obtaining a top-down view (Bird eye view) of the scene being imaged (MBLB91).

The inverse relationship corresponding to Equation (9.24), which transforms image points into world coordinates, can be written as:

\begin{displaymath}
\begin{bmatrix}
x_{i} \\
y_{i} \\
z_{i}
\end{bmatrix}=
\l...
...\mathbf{t}_0 = \lambda \mathbf{v}(u_{i}, v_{i}) + \mathbf{t}_0
\end{displaymath} (9.26)

where it is clear that each image point corresponds to a line in the world (as $\lambda$ varies) passing through the pin-hole ($\mathbf{t}_0$) and oriented in the direction
\begin{displaymath}
\mathbf{v}(u_i,v_i) = \mathbf{R}^{-1} \mathbf{K}^{-1}
\begin{bmatrix}
u_{i} \\
v_{i} \\
1
\end{bmatrix}\end{displaymath} (9.27)

where $\mathbf{v}: \mathbb{R}^2 \to \mathbb{R}^3$ is the function that associates each image point with the vector joining the pin-hole to the corresponding sensor point.

Using the Camera Matrix $\mathbf{P} = \left[ \mathbf{P}_{3 \times 3} \vert \mathbf{p}_4 \right]$ directly, it is possible to obtain a result equivalent to Equation (9.26) in the form

\begin{displaymath}
\begin{bmatrix}
x_{i} \\
y_{i} \\
z_{i}
\end{bmatrix}=
\l...
... \\
1
\end{bmatrix}-
\mathbf{P}_{ 3\times 3}^{-1}\mathbf{p}_4
\end{displaymath} (9.28)

without explicitly using the intrinsic and extrinsic parameter matrices. The two formulations are clearly equivalent.



Subsections
Paolo medici
2026-10-01