This chapter provides a general treatment of algorithms involving the analysis of images acquired by more than one camera, with particular emphasis on stereoscopic vision.
The ultimate goal of multicamera vision is not only to estimate the depth of observed points, but more generally to reconstruct a three-dimensional representation of the scene. Such a representation may take different forms, from traditional point clouds and meshes to more recent neural and differentiable representations.
These views may be temporally coincident (for example, in the case of a pair of cameras forming a stereo camera) or may observe the scene at different points in space and time, as occurs, for example, when processing images from the same camera as it moves through space (motion stereo, structure from motion).
Stereoscopic analysis can be implemented mainly through two techniques:
A necessary condition for obtaining a complete three-dimensional reconstruction of the observed scene through the analysis of multiple images acquired from different viewpoints is knowledge of the intrinsic parameters of the cameras involved and their relative pose.
If the relative pose is unknown, it can be estimated by analyzing the images themselves. However, as will be shown later, the distance between the cameras will be determined up to a multiplicative factor, and consequently the three-dimensional reconstruction will also be known only up to that factor.
Even if the intrinsic parameters are unknown, it is still possible to relate corresponding points in the two images and use this process to accelerate the matching of KeyPoints, but nothing can be said about the three-dimensional reconstruction of the observed scene (the reconstruction is known only up to an affine transformation).