Binary Classifiers

A particular and very common case of classifier is the binary classifier. In this case, the problem consists in finding a relationship linking the training set $S=\{ (\mathbf{x}_1, y_1) \ldots (\mathbf{x}_l, y_l) \} \in (\mathbb{X} \times \mathbb{Y})$ where $\mathbb{X} \subseteq \mathbb{R}^{n}$ is the vector collecting the information used for training and $\mathbb{Y}=\{+1,-1\}$ is the space of the associated classes.

Examples of binary classifiers are:

LDA
Linear Discriminant Analysis (section 5.3) is a technique for finding the separating plane between classes that maximizes the distance between the distributions;
Decision Stump
Single-level decision trees have only two possible outputs;
SVM
Support Vector Machines (section 5.5) partition the feature space, maximizing the margin, using hyperplanes or simple surfaces.

Linear classifiers (LDA and Linear SVM) are of particular interest because, in solving the binary classification problem, they identify a hyperplane $(\mathbf{w},b)$ separating the two classes.

The equation of a hyperplane, obtained by slightly modifying formula (1.85), is

\begin{displaymath}
\mathbf{w} \cdot \mathbf{x} + b = 0
\end{displaymath} (5.4)

where the normal vector $\mathbf{w}$ need not have unit norm. A hyperplane divides space into two subspaces in which equation (5.4) has opposite signs. The separating surface is a hyperplane that divides space into two regions representing the two categories of the binary classification.

A linear classifier is based on a discriminant function

\begin{displaymath}
f(\mathbf{x}) = \mathbf{w} \cdot \mathbf{x} + b
\end{displaymath} (5.5)

Vector $\mathbf{w}$ is called the weight vector, and term $b$ is called the bias. Linear classifiers are important because, by projecting along axis $\mathbf{w}$, they transform the problem from multidimensional to scalar.

The sign of function $f(\mathbf{x})$ represents the classification result. A separating hyperplane is equivalent to finding a linear combination of the elements $\mathbf{x} \in \mathbf{X}$ such that

\begin{displaymath}
\hat{y} = \sgn ( \mathbf{w} \cdot \mathbf{x} + b )
\end{displaymath} (5.6)

Paolo medici
2026-10-06