The RANdom Sample And Consesus algorithm is an iterative algorithm for estimating the parameters of a model when the data set is strongly affected by outliers; algorithm (FB81) is nondeterministic and based on the random selection of the elements that generate the model.
RANSAC and all its variants can be viewed as algorithms that iteratively alternate between two phases: hypothesis generation and hypothesis evaluation.
In brief, the algorithm consists of randomly selecting samples from all
input samples
, with
large enough to determine a model (the hypothesis).
Once a hypothesis has been obtained, the number of
elements of
that are sufficiently close to it to belong to it is counted.
A sample
belongs or does not belong to the hypothesized model (that is, it is a hypothesized inlier or outlier) according to whether its distance from model
is below or above a given threshold
, which normally depends on the problem.
The threshold
presents a difficulty in practical problems where the additive error is Gaussian, that is, where the support is infinite.
In this case, it is nevertheless necessary to define a probability
of detecting inliers in order to define a threshold
.
All samples satisfying the hypothesis are called consensus samples (consensus).
The set of consensus samples associated with hypothesis
is the consensus set of
:
| (4.132) |
Among all randomly generated models, the model satisfying a given metric is selected; in the original RANSAC, for example, this is the model with the largest consensus-set cardinality.
One problem is determining how many hypotheses to generate in order to have a high probability of obtaining the correct model.
There is a statistical relationship between the number of iterations and the probability
of finding a solution consisting entirely of inliers.
The number of trials
must satisfy
, that is,
| (4.133) |
Normally, can be approximated by
4.4, and therefore
| (4.134) |
Normally, is chosen to be equal to the number of elements required to create the model. However, if it is larger than this number, the generated model must be constructed through numerical regression with respect to the given constraints. This is necessary when the noise variance is high, although it increases the risk of including outliers among the elements satisfying the constraints.