Multi-Camera Vision

This chapter generally addresses algorithms involving the analysis of images acquired by more than one camera, with particular emphasis on the case of stereoscopic vision.

The ultimate goal of multi-camera vision is not merely to estimate the depth of the observed points, but more generally to reconstruct a three-dimensional representation of the scene. Such a representation can take different forms, from traditional point clouds and meshes to more recent neural and differentiable representations.

These views may be temporally coincident (for example, in the case of a pair of cameras forming a stereo camera) or may observe the scene from different points in space and time, as occurs, for example, when processing images from the same camera as it moves through space (motion stereo, structure from motion).

Stereo analysis can be implemented primarily through two techniques:

A necessary condition for obtaining a metric three-dimensional reconstruction of the observed scene from the analysis of multiple images acquired from different viewpoints is knowledge of the intrinsic parameters of the cameras involved and their relative pose.

If the relative pose is unknown, it can be estimated through image analysis itself. However, as will be shown below, the distance between the cameras will be recovered up to a multiplicative factor, and consequently the three-dimensional reconstruction will also be known only up to that factor.

Even if the intrinsic parameters are unknown, it is still possible to establish correspondences between points in the two images and, through this process, accelerate the matching of KeyPoints. In this case, the scene can be reconstructed only up to a projective transformation; obtaining a metric reconstruction requires additional calibration information.



Subsections
Paolo medici
2026-10-06