Ensemble Learning

The concept of Ensemble training involves using several different classifiers, combined in a certain way to maximize performance by exploiting the strengths of each classifier while limiting the weaknesses of the individual classifiers.

The concept of Ensemble Learning is based on weak classifiers (weak classifier): a weak classifier can correctly classify at least $50\%+1$ of the samples in a binary problem. Combined in a certain way, weak classifiers make it possible to construct a strong classifier while simultaneously addressing problems typical of traditional classifiers, especially overfitting.

The origins of Ensemble Learning, the concept of a weak classifier, and, in particular, the concept of probably approximately correct learning (PAC) were first introduced by Valiant (Val84).

In practice, Ensemble Learning techniques do not provide general-purpose classifiers; rather, they indicate the optimal way to combine multiple classifiers.

Examples of Ensemble Learning techniques include

Decision Tree
decision trees, being constructed from many Decision Stump classifiers arranged in cascade, are an early example of Ensemble Learning;
Bagging
BootStrap AGGregatING attempts to reduce overfitting by training different classifiers on subsets of the training set and then performing a majority vote;
Boosting
rather than selecting purely random subsets of the training set, some of the samples that remain incorrectly classified are used;
AdaBoost
ADAptive BOOSTing (Section 5.8.2) is the best-known Ensemble Learning algorithm and the progenitor of the large family of AnyBoost classifiers;
Random ForestTM
is a BootStrap Aggregating (bagging) of Decision Trees, an Ensemble Classifier composed of several decision trees, each created from a subset of the training data and features to be analyzed, which vote by majority;
and many others.

Examples of weak classifiers widely used in the literature are Decision Stumps (AL92) associated with Haar features (Section 7.1). The Decision Stump is a binary classifier of the form

\begin{displaymath}
h(\mathbf{x}) = \left\{ \begin{array}{ll}
+1 & \quad \text...
...theta \\
-1 & \quad \text{otherwise} \\
\end{array}\right.
\end{displaymath} (5.72)

where $f(\mathbf{x})$ is a function that extracts a scalar from the sample to be classified, $p=\{ +1, -1 \}$ is a parity indicating the direction of the inequality, and $\theta $ is the decision threshold (Figure 5.4).



Subsections
Paolo medici
2026-10-01