Regression and Optimization Methods for Model Analysis

One of the most widespread problems in computer vision (and, more generally, in information theory) is fitting a set of noisy measurements (for example, the pixels of an image) to a predefined model.

In addition to noise, which may be white Gaussian noise but potentially follow any statistical distribution, one must also consider the possible presence of outliers, a term used in statistics to denote data points that are too far from the model to actually belong to it.

This chapter presents both several regression techniques for estimating the parameters $\boldsymbol\beta$ of a stationary model from a set of noisy data and techniques for detecting and removing outliers from the input data.

The next chapter, instead, presents “regression” techniques more closely related to classification.

Some techniques found in the literature for estimating model parameters are the following:

Least Squares Fitting
If all data points are inliers, there are no outliers, and the only disturbance is additive white Gaussian noise, least-squares regression is the optimal technique (Section 4.2);
M-Estimator
Even a few outliers can significantly shift the model because the errors are squared (Hub96): assigning a nonquadratic weight to points far from the estimated model improves the estimate itself (Section 4.8);
IRLS
iteratively reweighted least squares is used when the outliers are very far from the model and relatively few in number: under these conditions, iterative regression can be performed (Section 4.9), where at each iteration points with excessively large errors are removed (ILS) or assigned different weights (IRLS);
Hough
If the input data are affected both by error and by many outliers, and a multimodal distribution may be present, but the model consists of few parameters, the Hough transform (Hou59) makes it possible to obtain the statistically most prevalent model (Section 4.11);
RANSAC
If the outliers are comparable in number to the inliers and the noise is very small (relative to the positions of the outliers), RANdom SAmpling and Consensus (FB87) makes it possible to obtain the best model present in the scene (Section 4.12);
LMedS
Least Median of Squares is an algorithm, similar to RANSAC, that orders the points according to their distance from the randomly generated model and selects, among all models, the one with the smallest error median (Rou84) (Section 4.12.2);
Kalman
Finally, a Kalman filter can be used to estimate the parameters of a model (see 3.9) when this information is required at run time.

Only RANSAC and the Hough transform can optimally handle the case in which two or more distributions simultaneously approach the model in the measurements.

There is also nothing to prevent the use of mixed techniques, such as a sufficiently coarse Hough transform (and therefore a fast method with low memory requirements) to remove the outliers, followed by least-squares regression to obtain the model parameters more accurately.



Subsections
Paolo medici
2026-10-06