Regression and Optimization Methods for Model Analysis

One of the most widespread problems in computer vision (and, more generally, in information theory) is fitting a set of noisy measurements (for example, the pixels of an image) to a predefined model.

In addition to noise, which may be white Gaussian noise but potentially follow any statistical distribution, the possible presence of outliers must also be considered. This term is used in statistics to denote data points that are too far from the model to actually belong to it.

This chapter presents both various regression techniques for estimating the parameters $\boldsymbol\beta$ of a stationary model from a set of noisy data and techniques for detecting and removing outliers from the input data.

The next chapter will instead present “regression” techniques more closely related to classification.

Some techniques found in the literature for estimating model parameters are the following:

Least Squares Fitting
If all the data are inliers, there are no outliers, and the only disturbance is additive white Gaussian noise, least-squares regression is the optimal technique (Section 4.2);
M-Estimator
Even the presence of a few outliers can significantly shift the model because errors are weighted quadratically (Hub96): assigning nonquadratic weights to points far from the estimated model improves the estimate itself (Section 4.8);
IRLS
iteratively reweighted least squares is used when the outliers are very far from the model and small in number. Under these conditions, iterative regression can be performed (Section 4.9), in which, at each iteration, points with excessively large errors are removed (ILS) or assigned different weights (IRLS);
Hough
If the input data are affected both by error and by many outliers, and a multimodal distribution may be present, but the model is described by only a few parameters, the Hough transform (Hou59) makes it possible to obtain the statistically most prevalent model (Section 4.11);
RANSAC
If the outliers are comparable in number to the inliers and the noise is very small (relative to the positions of the outliers), RANdom SAmpling and Consensus (FB87) makes it possible to obtain the best model present in the scene (Section 4.12);
LMedS
Least Median of Squares is an algorithm, similar to RANSAC, that orders the points according to their distance from the randomly generated model and selects the model with the smallest error median (Rou84) (Section 4.12.2);
Kalman
Finally, a Kalman filter can be used to estimate the parameters of a model (see 3.9) when this information is required at run time.

Only RANSAC and the Hough transform can handle optimally the case in which two or more distributions present in the measurement simultaneously approach the model.

There is also nothing to prevent the use of hybrid techniques, such as a sufficiently coarse Hough transform (and therefore a fast method with low memory requirements) for removing the outliers, followed by least-squares regression to obtain the model parameters more accurately.



Subsections
Paolo medici
2026-10-01