Even in the 3D–2D approach, the three-dimensional points are not known exactly but have been obtained through a previous triangulation procedure. Treating them as error-free data therefore artificially separates structure estimation from motion estimation.
A more comprehensive approach is to treat both the camera pose and the three-dimensional point positions as unknowns and to directly minimize the reprojection errors in the images. Under the assumption of independent, isotropic Gaussian errors in the image coordinates, this minimization corresponds to maximum-likelihood estimation:
The estimated projections can be expressed as
to which the constraint given by equation (10.99) must be added, while retaining the actual three-dimensional positions of the points
and
at the two instants in time as unknowns.
In this way, both camera motion and the three-dimensional position of each feature are estimated simultaneously. Maximum-likelihood estimation again requires solving a nonlinear problem, but with a number of unknowns that grows with the number of points considered. For a rectified stereo pair, the cost function can be further simplified by exploiting the particular geometry of the configuration.