The methods considered so far leave considerable freedom in choosing the loss function.
In practical cases where the cost function is quadratic, further optimizations over Newton's method can be introduced, avoiding the costly computation of the Hessian.
In this case, the loss function takes the form introduced earlier:
With this cost function, the gradient and Hessian are written as:
| (4.47) |
When the parameters are close to the solution, the residual is small and the Hessian can be approximated by the first term:
| (4.48) |
Under these conditions, the gradient and Hessian of the cost function depend only on the Jacobian of the functions
.
The Hessian approximated in this way can be substituted into equation (4.35):
As in Newton's method, this yields a linear minimization problem that can be solved using the normal equations:
The meaning of the normal equations is geometric: the minimum is reached when
is orthogonal to the column space of
.
In the particular case where the residual is written as:
| (4.51) |