Given a calibrated camera and a set of correspondences between three-dimensional world points
and their corresponding image observations
, the problem consists in determining the camera pose, namely the rotation
and translation
that best explain the available observations.
This problem is known in the literature as Perspective-n-Point (PnP), where denotes the number of three-dimensional correspondences used. In the minimum theoretical case of three points, the problem is called Perspective-3-Point (P3P). In the ideal case of a calibrated camera, three independent 3D–2D correspondences provide the minimum number of constraints required to determine the sensor pose.
One possible solution is to formulate the problem as a linear system and solve it using DLT (Direct Linear Transform) techniques. Although conceptually simple, this approach is generally inaccurate and numerically unstable when the observations are noisy.
A more natural approach is to observe that a candidate pose
makes it possible to project each three-dimensional point onto the image plane. The quality of the solution can therefore be measured using the reprojection error, defined as
| (9.76) |
where
denotes the projection of the three-dimensional point
obtained using the current pose.
The PnP problem can therefore be expressed as
| (9.77) |
that is, as the search for the pose that minimizes the total reprojection error.
Iterative methods that directly minimize this cost function generally provide highly accurate results, but require an initial estimate sufficiently close to the correct solution to avoid convergence to undesirable local minima.
For this reason, numerous algorithms specifically designed for the PnP problem have been developed. These algorithms can provide a robust and efficient initial pose estimate. The best-known include EPnP (LFNP09), which introduces a particularly efficient linear formulation of the problem, and DLS (Direct Least Squares) (HR11), which formulates pose estimation as a direct least-squares minimization.
In practical applications, pose estimation is generally performed by combining a PnP algorithm with robust techniques for removing outliers, such as RANSAC, followed by a nonlinear refinement stage based on minimizing the reprojection error.
The PnP problem plays a central role in numerous computer-vision applications, including visual localization, augmented reality, visual odometry, Structure from Motion, and SLAM. The main solution techniques will be examined in greater detail in the chapters devoted to multicamera vision and pose estimation.
Paolo medici