This chapter addresses the problem of describing the process through which the light incident on objects is recorded by a digital sensor. This concept is fundamental in image processing, as it provides the relationship linking image points to their position in the world; in other words, it makes it possible to determine the region of the world associated with an image pixel or, conversely, to identify the area of the image that captures a given region in world coordinates.
The universally accepted projective model, known as the Pin-Hole Camera model, is based on simple geometric relationships9.1.
Figure 9.1 shows a highly simplified diagram of how the image is formed on the sensor.
The observed point
, expressed in camera coordinates, is projected onto the sensor plane.
All these rays pass through the same point: the projection center (pin-hole).
|
Analyzing Figure 9.1, we observe that the ratios between the similar triangles formed by the optical rays yield the equation that projects a generic point
, expressed in camera coordinates (one of the reference frames in which operations can be performed), onto the sensor plane:
It should be noted that the coordinates
, expressed in camera coordinates, follow the left-hand rule in this book (widely used in computer graphics), as opposed to the right-hand rule (more commonly used in robotic applications), which is instead used for expressing world coordinates.
The coordinate
therefore represents the point's depth along the camera's optical axis.
Defining
Equation (9.1) can therefore also be written as
| (9.3) |
Normalized coordinates thus constitute an intermediate space between the three-dimensional geometry of the scene and the discrete image coordinates. Many computer vision algorithms, such as pose estimation, triangulation, and epipolar geometry, assume that image points are expressed in this reference frame.
The sensor coordinates
are not yet the coordinates of the digital image, but rather “intermediate” coordinates expressed in the physical units of the sensor.
An additional transformation is therefore required to obtain the image coordinates:
and
are conversion factors between the units of the sensor reference frame, typically expressed in metric units, and those of the image, expressed in pixels.
For square pixels, the two conversion factors are equal.
In the absence of information (available in the various datasheets) about ,
, and
, these variables are often combined into two new parameters called
and
, namely the effective focal lengths measured in pixels, which can be obtained empirically from images, as discussed in the section on calibration.
These variables are defined as
Using normalized image coordinates, the transformation to pixel coordinates therefore becomes
When the principal point approximately coincides with the image center, and
can also be related to half the field of view:
| (9.7) |
When the sensor has square pixels, and
tend to have the same value.
Because of the division by , Equation (9.1) cannot be represented directly by a linear transformation in Cartesian coordinates.
However, it is possible to reformulate it by introducing a scale factor
, thereby representing the projection through a linear system in homogeneous coordinates.
To do so, we will use the theory presented in Section 1.5 concerning homogeneous coordinates.
The transformation between normalized image coordinates and pixel coordinates can be represented by the intrinsic-parameter matrix
When , we therefore have
| (9.9) |
Recalling that
For this reason, is normally omitted and homogeneous coordinates are used: to obtain the point in non-homogeneous coordinates, it is sufficient to divide the first two coordinates by the third.
The use of homogeneous coordinates therefore makes the division by coordinate
implicit.
Matrix depends exclusively on the camera's internal parameters and is therefore called the intrinsic-parameter matrix.
In the general case, it is an upper-triangular matrix defined by five parameters.
With modern digital sensors, it is normally possible to set the skew factor , which accounts for the possibility that the angle between the sensor axes may not be exactly
, to zero.
Setting , the inverse of matrix (9.8) can be written as
In particular, the inverse intrinsic matrix makes it possible to transform an image point into the corresponding normalized image coordinates:
| (9.12) |
Knowledge of these parameters (see Section 9.5 on calibration) therefore makes it possible to transform a point from camera coordinates to image coordinates or, conversely, to generate the line in camera coordinates associated with an image point.
In particular, the vector
| (9.13) |
This model, however, does not account for the effects of lens distortion. The pin-hole camera model is valid only when the image coordinates used refer to distortion-free images.