Also in the 3D–2D approach, the three-dimensional points are not known exactly, but have been obtained through a previous triangulation procedure. Treating them as error-free data therefore artificially separates structure estimation from motion estimation.
A more complete approach consists in treating both the camera pose and the three-dimensional positions of the points as unknowns, and directly minimizing the reprojection errors in the images. Under the assumption of independent and isotropic Gaussian errors in the image coordinates, this minimization corresponds to the maximum-likelihood estimate:
The estimated projections can be expressed as
with the constraint given by equation (10.102), while also treating the actual three-dimensional positions of the points
and
at the two instants as unknowns.
In this way, both camera motion and the three-dimensional positions of the individual features are estimated simultaneously. The maximum-likelihood solution again requires solving a nonlinear problem, but with a number of unknowns that grows with the number of points considered. In the case of a rectified stereo pair, the cost function can be further simplified by exploiting the particular geometry of the configuration.