Model Parameter Evaluation

Ignoring the presence of outliers in the input data used for regression, two important open questions remain: assessing how good the resulting model is and, at the same time, providing an indication of how far this estimate is from the true model because of errors in the input data.

This section discusses the nonlinear case in detail: the linear case is equivalent, using the parameter matrix $\mathbf{X}$ in place of the Jacobian $\mathbf{J}$, and was partly addressed in Section 1.2.

Let $\mathbf{y}=\left(y_1, \ldots, y_n \right)^{\top}$ be a vector of realizations of statistically independent random variables $y \in \mathbb{R}$ and $\boldsymbol\beta \in \mathbb{R}^m$ model parameters. An intuitive estimator of model quality is the root-mean-squared residual error (RMSE), also called the standard error of the regression:

\begin{displaymath}
s = \sqrt{ \frac{ \sum^{n}_{i=1} \left( y_{i} - \hat{y}_{i} \right)^{2} } {n} }
\end{displaymath} (4.84)

where $\hat{y}_i = f(\mathbf{x}_i, \hat{\boldsymbol\beta} )$ is the value estimated using the model $f$ from which the parameters $\hat{\boldsymbol\beta}$ were obtained. This function has usually already been encountered in residual form, $r_i = \mathbf{y}_i - \hat{\mathbf{y}}_i$. If the estimator is unbiased (as occurs, for example, in least-squares regression), $\E [ r_i ] = 0$. Therefore, when the observation noise is zero-mean Gaussian, the value of $s \geq \sigma$, and the two values are equal when the model is optimal.

This, however, is not a direct measure of the quality of the solution found; it only measures how closely the model matches the input data. Consider, for example, the limiting case of systems that are not overdetermined, where the residual is always zero, regardless of the amount of noise affecting the individual observations.

The most suitable measure for evaluating the model is the parameter variance-covariance matrix (Parameter Variances and Covariances matrix).

Covariance forward propagation (covariance forward propagation) was already presented in Section 2.6. Briefly, there are three methods for performing this operation: the first is based on a linear approximation of the model and involves the Jacobian; the second is based on the more general Monte Carlo simulation technique; and, finally, a modern alternative midway between the first two is the Unscent Transformation (Section 3.5), which empirically provides estimates up to third order in the case of Gaussian noise.

Evaluating the quality of the estimated parameters $\hat{\boldsymbol\beta}$ given the estimated noise covariance (Covariance Matrix Estimation) is exactly the opposite problem, because it requires computing the backward propagation of the variance (backward propagation). Indeed, once this covariance matrix has been obtained, it is possible to define a confidence interval around $\hat{\boldsymbol\beta}$.

The quality of the parameter estimate $\hat{\boldsymbol\beta}$ in the nonlinear case can be evaluated to first order by inverting the linearized version of the model (although, here too, techniques such as Monte Carlo or the UT can be used for more rigorous estimates).

The covariance matrix associated with the proposed solution $\hat{\boldsymbol\beta}$ can be determined when the function $f$ is bijective and differentiable in a neighborhood of that solution. Let $f : \mathbb{R}^m \to \mathbb{R}^n$ be a multivariate multidimensional function. It is possible to estimate the mean $\bar{\mathbf{r}} = \E \left[\mathbf{y} - f(\hat{\boldsymbol\beta}) \right] \approx \mathbf{0}$ and the cross-covariance matrix $\boldsymbol\Sigma_r$ of the residuals; then the inverse transformation $f^{-1}$ will have mean $\hat{\boldsymbol\beta}$ and covariance matrix

\begin{displaymath}
\Sigma_{\boldsymbol\beta} = (\mathbf{J}^{\top} \Sigma_r^{-1} \mathbf{J})^{-1}
\end{displaymath} (4.85)

where $\mathbf{J}$ is the Jacobian of the model $f$ evaluated at point $\hat{\boldsymbol\beta}$:
\begin{displaymath}
J_{i,j} = \frac{\partial r_i}{\partial \beta_j } (\hat{\bold...
...frac{\partial f_i}{\partial \beta_j } (\hat{\boldsymbol\beta})
\end{displaymath} (4.86)

Equation (4.85) is obtained by manipulating equation (2.38), which computes the forward propagation of uncertainty.

Note that this (the inverse of the information matrix) is the Cramér–Rao lower bound on the covariance that an unbiased estimator of the parameter $\boldsymbol\beta$ can have.

When the transformation $f$ is underdetermined, the rank of the Jacobian $d$, with $d<m$, is called the number of essential parameters (essential parameters). For an underdetermined transformation $f$, formula (4.85) is not invertible, but it can be shown that the best approximation of the covariance matrix can be obtained using the pseudoinverse:

\begin{displaymath}
\Sigma_{\boldsymbol\beta} = (\mathbf{J}^{\top} \Sigma_r^{-1} \mathbf{J})^{+}
\end{displaymath}

Alternatively, one can perform a QR decomposition with column pivoting of the Jacobian, identify the linearly dependent columns (by analyzing the diagonal of matrix R), and remove them during the matrix inversion itself. In the very common case in which $f$ is a scalar function and the observation noise is independent and has constant variance, the asymptotic covariance matrix (Asymptotic Covariance Matrix) can be written more simply as
\begin{displaymath}
\Sigma_{\boldsymbol\beta} = ( \mathbf{J}^{\top}\mathbf{J})^{-1} \sigma^{2}
\end{displaymath} (4.87)

where $\sigma^2$ is the observation-noise variance, under assumption $\boldsymbol \Sigma_r = \sigma^2 \mathbf{I}$, which is valid for independent realizations. Since $\mathbf{J}$ depends only on the geometry of the problem, the matrix $( \mathbf{J}^{\top}\mathbf{J})^{-1}$ also depends only on the problem and not on the observations. Asymptotically, the estimate tends to $\boldsymbol\beta = \mathcal{N} \left( \hat{\boldsymbol\beta}, \Sigma_{\boldsymbol\beta} \right)$. Since the Jacobian matrix indicates how sensitive the outputs are to the parameters, it is also called the sensitivity matrix.

The observation-noise estimate may be empirical, assuming by the law of large numbers that $\sigma = s$, and computed as

\begin{displaymath}
\sigma^{2} \approx \frac{\sum_{i=1}^{n} r_i^2}{n-m}
\end{displaymath} (4.88)

using the posterior statistics of the error in the data $r_i$. The denominator $n-m$ represents the statistical degrees of freedom of the problem: thus, the estimated variance is infinite when the number of unknown model parameters equals the number of data points collected.

The Eicker–White covariance estimator is slightly different and is left to the reader for further study.

The parameter variance-covariance matrix represents the error ellipsoid.

A useful metric for assessing the problem is the D-optimal design:

\begin{displaymath}
\det \left( \mathbf{J}^{\top}\mathbf{J} \right)^{-1}
\end{displaymath} (4.89)

which minimizes the determinant of the variance-covariance matrix or, equivalently, maximizes the Fisher information matrix:
\begin{displaymath}
\det \mathbf{F}\left( \boldsymbol\beta \right)
\end{displaymath} (4.90)

Geometrically, this approach minimizes the volume of the error ellipsoid.

Other metrics include, for example, the E-optimal design, which consists in maximizing the smallest eigenvalue of the Fisher matrix, or equivalently minimizing the largest eigenvalue of the variance-covariance matrix. Geometrically, this minimizes the maximum diameter of the ellipsoid.

Paolo medici
2026-10-06