Binary Classifiers

A particular and very common type of classifier is the binary classifier. In this case, the problem consists in finding a relationship linking the training set $S=\{ (\mathbf{x}_1, y_1) \ldots (\mathbf{x}_l, y_l) \} \in (\mathbb{X} \times \mathbb{Y})$ where $\mathbb{X} \subseteq \mathbb{R}^{n}$ is the vector collecting the information to be used for training and $\mathbb{Y}=\{+1,-1\}$ is the space of the associated classes.

Examples of intrinsically binary classifiers include:

LDA
Linear Discriminant Analysis (section 5.3) is a technique for finding the separating plane between classes that maximizes the distance between the distributions;
Decision Stump
Single-level decision trees have only two possible outputs;
SVM
Support Vector Machines (Support Vector Machines, section 5.5) partition the feature space, maximizing the margin, using hyperplanes or simple surfaces.

Linear classifiers (LDA and Linear SVM) are of particular interest because, in solving the binary classification problem, they identify a hyperplane $(\mathbf{w},b)$ separating the two classes.

The equation of a hyperplane, obtained by slightly modifying formula (1.85), is

\begin{displaymath}
\mathbf{w} \cdot \mathbf{x} + b = 0
\end{displaymath} (5.4)

where the normal vector $\mathbf{w}$ need not have unit norm. A hyperplane divides space into two subspaces in which equation (5.4) has opposite signs. The separating surface is a hyperplane that divides space into two subregions representing the two categories of the binary classification.

A linear classifier is based on a discriminant function

\begin{displaymath}
f(\mathbf{x}) = \mathbf{w} \cdot \mathbf{x} + b
\end{displaymath} (5.5)

The vector $\mathbf{w}$ is called the weight vector, and the term $b$ is called the bias. Linear classifiers are important because, through projection along the $\mathbf{w}$ axis, they transform the problem from multidimensional to scalar.

The sign of function $f(\mathbf{x})$ represents the classification result. A separating hyperplane is equivalent to identifying a linear combination of the elements $\mathbf{x} \in \mathbf{X}$ such that

\begin{displaymath}
\hat{y} = \sgn ( \mathbf{w} \cdot \mathbf{x} + b )
\end{displaymath} (5.6)

Paolo medici
2026-10-01