Loss Functions and Risk Minimization

A general way to interpret many classification algorithms is to view them as the result of minimizing an appropriate loss function (loss function).

For the linear classifiers introduced in the previous sections, classification is generally obtained from a discriminant function of the form


\begin{displaymath}
f(\mathbf{x})
=
\mathbf{w}^{\top}\mathbf{x}+b.
\end{displaymath} (5.50)

Logistic regression (Section 5.4), for example, interprets this function through a logistic transformation, whereas the SVM (Section 5.5) directly uses the value of the discriminant function to determine the class and the separating margin.

Consider a binary classifier and assume that the labels are


\begin{displaymath}
y\in\{-1,+1\}.
\end{displaymath} (5.51)

The quantity


\begin{displaymath}
m = y f(\mathbf{x})
\end{displaymath} (5.52)

is called the functional margin. Its sign determines whether the classification is correct: $m>0$ corresponds to a correct classification, whereas $m<0$ identifies a misclassified sample.

The ideal objective of a classifier is to minimize the number of classification errors. This objective can be expressed by the loss function


\begin{displaymath}
L_{0/1}(m)
=
\begin{cases}
0 & m>0\\
1 & m\leq 0
\end{cases}\end{displaymath} (5.53)

known as the 0/1 loss.

Direct minimization of the 0/1 loss is computationally inconvenient, however. For this reason, many learning algorithms use surrogate loss functions, that is, functions that approximate the behavior of the 0/1 loss while being easier to optimize.



Subsections
Paolo medici
2026-10-01