Multiclass Classification

The classifiers introduced so far have mainly been presented in the binary setting. In practice, however, many classification problems require simultaneous discrimination among multiple categories. It is therefore necessary to extend the formalism of loss functions and linear classifiers to the multiclass case.

In binary classification, considered in the preceding sections, a linear classifier returns a single value

\begin{displaymath}
f(\mathbf{x}) = \mathbf{w}^{\top}\mathbf{x} + b
\end{displaymath} (5.63)

that determines on which side of the decision boundary the sample lies. In the multiclass case, with $K$ classes, it is instead necessary to associate a score with each class, yielding a vector
\begin{displaymath}
\mathbf{s} = f(\mathbf{x},\mathbf{W},\mathbf{b})
\end{displaymath} (5.64)

with
\begin{displaymath}
s_j = \mathbf{w}_j^{\top}\mathbf{x} + b_j,
\qquad j=1,\ldots,K.
\end{displaymath} (5.65)

The scores $s_j$ do not necessarily represent probabilities, and their absolute values generally have no direct meaning. The classification problem therefore consists in comparing the scores of the different classes and assigning the sample to the class with the highest score.

In the multiclass case as well, it is possible to define a loss function that quantifies how compatible the scores produced by the classifier are with the correct class. As discussed in Section 5.6, the choice of loss function substantially determines the criterion used to train the classifier.

One possibility is to extend to the multiclass case the concept of the hinge loss introduced for the SVM in Section 5.5. Let $y_i$ denote the correct class of the $i$-th sample and $s_j=f_j(\mathbf{x}_i)$ the score associated with class $j$. The multiclass SVM loss can then be defined as

\begin{displaymath}
\ell_i =
\sum_{j \neq y_i}
\max\left(0, s_j - s_{y_i} + 1\right).
\end{displaymath} (5.66)

The loss is zero when the score of the correct class exceeds that of every other class by at least a margin of $1$. Otherwise, every class that violates this margin is penalized. This formulation is therefore a natural extension of the binary hinge loss, in which the margin is evaluated with respect to all competing classes.

A variant is the squared hinge loss, in which the penalty is squared:

\begin{displaymath}
\ell_i =
\sum_{j \neq y_i}
\max\left(0, s_j - s_{y_i} + 1\right)^2.
\end{displaymath} (5.67)

The overall loss function can therefore be obtained as the average of the losses over the samples, optionally with an additional regularization term:

\begin{displaymath}
\mathcal{L} =
\frac{1}{n}\sum_{i=1}^{n}\ell_i
+
\lambda R(\mathbf{W}).
\end{displaymath} (5.68)

This formulation falls within the framework of empirical risk minimization introduced in Section 5.6. The term $R(\mathbf{W})$ penalizes excessively complex models and may, for example, correspond to L2 regularization of the weights.

Cross-entropy can naturally be extended to the multiclass case by introducing a probabilistic parameterization of the classes through the Softmax function.

The scores $s_j$ are transformed into probabilities through

\begin{displaymath}
p_j =
\frac{e^{s_j}}
{\sum_{k=1}^{K} e^{s_k}}.
\end{displaymath} (5.69)

The loss associated with the $i$-th sample is therefore the cross-entropy of the correct class:

\begin{displaymath}
\ell_i
=
-\log p_{y_i}
=
-\log
\frac{e^{s_{y_i}}}
{\sum_{j=1}^{K} e^{s_j}}
=
-s_{y_i}
+
\log\sum_{j=1}^{K}e^{s_j}.
\end{displaymath} (5.70)

This formulation is the multiclass extension of the cross-entropy introduced in Section 5.6. Since minimizing cross-entropy is equivalent to minimizing the negative log-likelihood of the observed class, the model can be interpreted as a maximum-likelihood estimator.

Here too, a regularization term can be introduced:

\begin{displaymath}
\mathcal{L} =
-\frac{1}{n}\sum_{i=1}^{n}
\log p_{y_i}
+
\lambda R(\mathbf{W}).
\end{displaymath} (5.71)

From a statistical perspective, the regularization term can be interpreted as introducing prior information about the model parameters. Under this interpretation, the estimate is no longer a simple maximum-likelihood estimate, but a Maximum A Posteriori (MAP) estimate.

Loss functions therefore provide a common language for describing numerous classification algorithms. An alternative approach consists in combining multiple simple classifiers to construct a more complex classifier, following the Ensemble Learning paradigm.

Paolo medici
2026-10-06