Model Parameter Evaluation

Ignoring the presence of outliers in the input data on which regression is performed, two important open issues remain: assessing how good the resulting model is and, at the same time, providing an indication of how far this estimate is from the true model because of errors in the input data.

This section discusses the nonlinear case in detail; the linear case is equivalent if the Jacobian $\mathbf{J}$ is replaced by the parameter matrix $\mathbf{X}$, as already partly discussed in Section 1.2.

Let $\mathbf{y}=\left(y_1, \ldots, y_n \right)^{\top}$ be a vector of realizations of statistically independent random variables $y \in \mathbb{R}$ and $\boldsymbol\beta \in \mathbb{R}^m$ model parameters. An intuitive estimator of model quality is the root-mean-squared residual error (RMSE), also called the standard error of the regression:

\begin{displaymath}
s = \sqrt{ \frac{ \sum^{n}_{i=1} \left( y_{i} - \hat{y}_{i} \right)^{2} } {n} }
\end{displaymath} (4.79)

where $\hat{y}_i = f(\mathbf{x}_i, \hat{\boldsymbol\beta} )$ is the value estimated by the model $f$ from which the parameters $\hat{\boldsymbol\beta}$ were obtained. This function is normally already familiar in its residual form $r_i = \mathbf{y}_i - \hat{\mathbf{y}}_i$. If the estimator is unbiased (as occurs, for example, in least-squares regression), $\E [ r_i ] = 0$. Therefore, when the observation noise is zero-mean Gaussian, the value of $s \geq \sigma$, and the two values are equal when the model is optimal.

However, this is not a direct measure of the quality of the solution found, but only of how well the model matches the input data: consider, for example, the limiting case of systems that are not overdetermined, where the residual is always zero, regardless of the amount of noise affecting the individual observations.

The most appropriate measure for evaluating the model is the parameter variance-covariance matrix (Parameter Variances and Covariances matrix).

Covariance forward propagation (covariance forward propagation) was already presented in Section 2.6 and, by way of a brief reminder, there are three methods for carrying out this operation: the first is based on a linear approximation of the model and involves the use of the Jacobian; the second is based on the more general Monte Carlo simulation technique; and finally, a modern alternative midway between the first two is the Unscent Transformation (Section 3.5), which empirically provides estimates up to third order in the case of Gaussian noise.

Evaluating the quality of the estimated parameters $\hat{\boldsymbol\beta}$ given the estimated noise covariance (Covariance Matrix Estimation) is exactly the opposite problem, because it requires computing the backward propagation of the variance (backward propagation). Once this covariance matrix has been obtained, it is possible to define a confidence interval in the neighborhood of $\hat{\boldsymbol\beta}$.

The quality of the parameter estimate $\hat{\boldsymbol\beta}$ in the nonlinear case can be evaluated to first order by inverting the linearized version of the model (although in this case techniques such as Monte Carlo or the UT can be used for more rigorous estimates).

It is possible to determine the covariance matrix associated with the proposed solution $\hat{\boldsymbol\beta}$ when the function $f$ is one-to-one and differentiable in the neighborhood of that solution. Let $f : \mathbb{R}^m \to \mathbb{R}^n$ be a multivariate multidimensional function. It is possible to estimate the mean $\bar{\mathbf{r}} = \E \left[\mathbf{y} - f(\hat{\boldsymbol\beta}) \right] \approx \mathbf{0}$ and the cross-covariance matrix $\boldsymbol\Sigma_r$ of the residuals; then the inverse transformation $f^{-1}$ will have mean $\hat{\boldsymbol\beta}$ and covariance matrix

\begin{displaymath}
\Sigma_{\boldsymbol\beta} = (\mathbf{J}^{\top} \Sigma_r^{-1} \mathbf{J})^{-1}
\end{displaymath} (4.80)

where $\mathbf{J}$ is the Jacobian of the model $f$ evaluated at the point $\hat{\boldsymbol\beta}$:
\begin{displaymath}
J_{i,j} = \frac{\partial r_i}{\partial \beta_j } (\hat{\bold...
...frac{\partial f_i}{\partial \beta_j } (\hat{\boldsymbol\beta})
\end{displaymath} (4.81)

Equation (4.80) is obtained by manipulating equation (2.38), which computes the forward propagation of uncertainty.

It should be noted that this quantity (the inverse of the information matrix) is the Cramér-Rao lower bound on the covariance that an unbiased estimator of the parameter $\boldsymbol\beta$ can have.

When the transformation $f$ is underdetermined, the rank of the Jacobian $d$, with $d<m$, is called the number of essential parameters (essential parameters). For an underdetermined transformation $f$, formula (4.80) cannot be inverted, but the best approximation of the covariance matrix can be obtained using the pseudoinverse:

\begin{displaymath}
\Sigma_{\boldsymbol\beta} = (\mathbf{J}^{\top} \Sigma_r^{-1} \mathbf{J})^{+}
\end{displaymath}

Alternatively, a QR decomposition with column pivoting can be applied to the Jacobian, the linearly dependent columns can be identified (by analyzing the diagonal of matrix R), and they can be removed during the matrix inversion itself. In the very common case in which $f$ is a scalar function and the observation noise is independent with constant variance, the estimated asymptotic covariance matrix (Asymptotic Covariance Matrix) can be written more simply as
\begin{displaymath}
\Sigma_{\boldsymbol\beta} = ( \mathbf{J}^{\top}\mathbf{J})^{-1} \sigma^{2}
\end{displaymath} (4.82)

where $\sigma^2$ is the observation-noise variance, under assumption $\boldsymbol \Sigma_r = \sigma^2 \mathbf{I}$, which is valid for independent realizations. Since $\mathbf{J}$ depends only on the geometry of the problem, matrix $( \mathbf{J}^{\top}\mathbf{J})^{-1}$ also depends only on the problem and not on the observations. Asymptotically, the estimate tends to $\boldsymbol\beta = \mathcal{N} \left( \hat{\boldsymbol\beta}, \Sigma_{\boldsymbol\beta} \right)$. The Jacobian matrix, because it indicates how sensitive the outputs are to the parameters, is also called the sensitivity matrix.

The observation-noise estimate may be empirical, assuming by the law of large numbers that $\sigma = s$, and computed using

\begin{displaymath}
\sigma^{2} \approx \frac{\sum_{i=1}^{n} r_i^2}{n-m}
\end{displaymath} (4.83)

using the posterior statistics of the error in the data $r_i$. The denominator $n-m$ represents the statistical degrees of freedom of the problem: thus, the estimated variance is infinite when the number of unknown model parameters equals the number of data points collected.

The Eicker-White covariance estimator is slightly different and is left for the reader to study.

The parameter variance-covariance matrix represents the error ellipsoid.

A useful metric for evaluating the problem is the D-optimal design:

\begin{displaymath}
\det \left( \mathbf{J}^{\top}\mathbf{J} \right)^{-1}
\end{displaymath} (4.84)

which minimizes the determinant of the variance-covariance matrix, or, equivalently, maximizes the Fisher information matrix:
\begin{displaymath}
\det \mathbf{F}\left( \boldsymbol\beta \right)
\end{displaymath} (4.85)

Geometrically, this approach minimizes the volume of the error ellipsoid.

Other metrics include, for example, the E-optimal design, which maximizes the smallest eigenvalue of the Fisher matrix, or equivalently minimizes the largest eigenvalue of the variance-covariance matrix. Geometrically, this minimizes the maximum diameter of the ellipsoid.

Paolo medici
2026-10-01