Main Loss Functions

Different classification techniques use different loss functions. The choice of loss function determines which characteristics of the errors are penalized most heavily during training.

Logistic regression (Section 5.4) uses the logistic loss


\begin{displaymath}
L_{\mathrm{log}}(m)
=
\log\left(1+e^{-m}\right).
\end{displaymath} (5.57)

The loss decreases as the margin increases and continues to penalize correctly classified samples as well. For negative values of $m$, corresponding to incorrect classifications, the penalty increases rapidly.

It is important to note that this loss coincides with the cost function introduced in the section on logistic regression, although it is written in a different form.

In logistic regression, the classes are encoded as $y \in \{0,1\}$ and the cost function is expressed as binary cross-entropy. Using instead the encoding $y\in\{-1,+1\}$ and introducing the margin $m=yf(\mathbf{x})$ yields exactly the logistic loss. Thus, this is the same loss function written using two different representations of the problem.

The Soft Margin SVM, introduced in Section 5.5, instead uses the hinge loss


\begin{displaymath}
L_{\mathrm{hinge}}(m)
=
\max(0,1-m).
\end{displaymath} (5.58)

In this case, samples with $m\geq1$ no longer contribute to the loss function. The function therefore explicitly introduces the characteristic margin of the SVM.

The primal formulation of the problem, introduced in equation (5.40), can be rewritten using the hinge loss as


\begin{displaymath}
\min_{\mathbf{w},b}
\frac{1}{2}\Vert\mathbf{w}\Vert^2
+
C
\s...
...=1}^{N}
L_{\mathrm{hinge}}
\left(
y_i f(\mathbf{x}_i)
\right).
\end{displaymath} (5.59)

This form makes the connection clear between the original SVM formulation, based on the slack variables $\xi_i$, and the formulation based on minimizing a loss function.

Another particularly important loss function in probabilistic classification is the cross-entropy. In the multiclass case, with $K$ categories, it can be written as


\begin{displaymath}
L_{\mathrm{CE}}
=
-\sum_{k=1}^{K}
y_k\log p_k,
\end{displaymath} (5.60)

where $y_k$ represents the target distribution and $p_k$ the probability assigned by the model to class $k$.

In the binary case, cross-entropy takes the form


\begin{displaymath}
L_{\mathrm{CE}}
=
-\left[
y\log p
+
(1-y)\log(1-p)
\right].
\end{displaymath} (5.61)

This is the same cost function obtained by maximizing the log-likelihood in logistic regression (Section 5.4).

Another loss function used in classification algorithms is the exponential loss


\begin{displaymath}
L_{\mathrm{exp}}(m)
=
e^{-m},
\end{displaymath} (5.62)

used, for example, by AdaBoost. This function heavily penalizes samples with a negative margin and continues to assign a decreasing weight to correctly classified samples.

The main loss functions considered can therefore be compared as follows:

Loss function Typical method Main characteristic
0/1 loss – classification error
Logistic loss Logistic Regression probabilistic classification
Hinge loss SVM margin maximization
Exponential loss AdaBoost error penalization
Cross-entropy Multiclass probabilistic models probabilistic classification

The loss-function formulation therefore makes it possible to bring many classification algorithms within a common framework. The loss function measures the error made by the model, while any regularization term limits its complexity. The differences between the various algorithms lie, among other things, in the choice of loss function, the form of the model, and the criterion used to control its complexity.

The risk-minimization framework is not limited to binary classification. The same ideas can naturally be extended to the multiclass case, as shown in the following section.

Paolo medici
2026-10-01