Linear Regression with a Polynomial Function

The method used to obtain linear regression for a line expressed in explicit form can be generalized to any polynomial function of the form:

\begin{displaymath}
y = \beta_0 + \beta_1 x + \beta_2 x^2 + \ldots + \beta_m x^m + \varepsilon
\end{displaymath} (4.102)

where $\beta_0 \ldots \beta_m$ are the parameters of the curve to be determined, obtained by minimizing the error function described in (4.7). The derivatives of a polynomial function are noteworthy:
\begin{displaymath}
\begin{array}{rl}
\frac{\partial S}{\partial \beta_j} & = ...
...+ \ldots + \beta_m \sum x_i^{j+m} - \sum y_i x_i^j
\end{array}\end{displaymath} (4.103)

Setting the gradient to zero therefore means solving the associated system:

\begin{displaymath}
\begin{bmatrix}
\sum 1 & \ldots & \sum x_i^{m} \\
\sum x...
... \sum y_i x_i \\
\vdots \\
\sum y_i x_i^m \\
\end{bmatrix}\end{displaymath} (4.104)

which is a symmetric matrix.

Alternatively, one can use the theory of the pseudoinverse (Section 1.1) and directly use equation (4.102) to construct an overdetermined linear system:

\begin{displaymath}
\begin{bmatrix}
1 & x_1 & \ldots & x_1^{m} \\
1 & x_2 & ...
...bmatrix}
y_1 \\
y_2 \\
\vdots \\
y_n \\
\end{bmatrix}
\end{displaymath} (4.105)

the Vandermonde matrix. The solution of this system yields the coefficients of the polynomial that minimizes the squared residuals. If the pseudoinverse is computed using the normal equations method, the resulting system is seen to be exactly the same as equation (4.104).

As will be seen elsewhere in this book, matrices such as the Vandermonde matrix, whose columns have different orders of magnitude, are ill-conditioned and require normalization to improve numerical stability.

Paolo medici
2026-10-06