Gauss-Newton

The methods considered so far leave considerable freedom in choosing the loss function. In practical cases where the cost function $\ell$ is quadratic, further optimizations over Newton's method can be introduced, avoiding the costly computation of the Hessian.

In this case, the loss function takes the form introduced earlier:

\begin{displaymath}
S(\boldsymbol\beta) = \frac{1}{2} \mathbf{r}^{\top} \mathbf{r} = \frac{1}{2} \sum_{i=1}^{n} r_i^2 (\boldsymbol\beta)
\end{displaymath} (4.46)

The term $1/2$ is used to obtain a more compact expression for the gradient.

With this cost function, the gradient and Hessian are written as:

\begin{displaymath}
\begin{array}{l}
\nabla S(\boldsymbol\beta) = \sum_{i=1}^{...
...athbf{J}_{r} + \sum_{i=1}^{n} r_i \mathbf{H}_{r_i}
\end{array}\end{displaymath} (4.47)

When the parameters are close to the solution, the residual is small and the Hessian can be approximated by the first term:

\begin{displaymath}
\mathbf{H}_S(\boldsymbol\beta) \approx \mathbf{J}_{r}^{\top}\mathbf{J}_{r}
\end{displaymath} (4.48)

Under these conditions, the gradient and Hessian of the cost function $S$ depend only on the Jacobian of the functions $r_i(\boldsymbol\beta)$. The Hessian approximated in this way can be substituted into equation (4.35):

\begin{displaymath}
- \mathbf{J}_{r}^{\top} \mathbf{r} = \mathbf{H}_{S} \boldsy...
...ox \mathbf{J}_{r}^{\top}\mathbf{J}_{r} \boldsymbol\delta_\beta
\end{displaymath} (4.49)

As in Newton's method, this yields a linear minimization problem that can be solved using the normal equations:

\begin{displaymath}
\boldsymbol\delta_\beta = - \left( \mathbf{J}_{r}^{\top}\mathbf{J}_{r} \right)^{-1} \mathbf{J}_{r}^{\top} \mathbf{r}
\end{displaymath} (4.50)

The meaning of the normal equations is geometric: the minimum is reached when $\mathbf{J}\boldsymbol\delta_\beta - \mathbf{r}$ is orthogonal to the column space of $\mathbf{J}$. In the particular case where the residual is written as:

\begin{displaymath}
r_i = y_i - f_i(\mathbf{x}_i ; \boldsymbol\beta)
\end{displaymath} (4.51)

that is, as in equation (4.6), it is possible to use $\mathbf{J}_f$, the Jacobian of $f$, instead of $\mathbf{J}_r$:
\begin{displaymath}
\boldsymbol\delta_\beta = \left( \mathbf{J}_{f}^{\top}\mathbf{J}_{f} \right)^{-1} \mathbf{J}_{f}^{\top} \mathbf{r}
\end{displaymath} (4.52)

having observed that the derivatives of $r_i$ and $f_i(\mathbf{x}_i)$ coincide up to sign.4.2



Footnotes

... sign.4.2
The derivatives coincide when a residual of the form $r_i = \hat{y}_i - y_i$ is chosen.
Paolo medici
2026-10-01