Ensemble Learning

The concept of Ensemble training involves using different classifiers combined in a specific way to maximize performance by exploiting the strengths of each classifier while limiting their individual weaknesses.

At the foundation of Ensemble Learning are weak classifiers (weak classifier): a weak classifier can correctly classify at least $50\%+1$ of the samples in a binary problem. Combined in a particular way, weak classifiers make it possible to construct a strong classifier while simultaneously addressing typical problems of traditional classifiers, foremost among them overfitting.

The origins of Ensemble Learning, the concept of the weak classifier, and, above all, the concept of probably approximately correct learning (PAC) are due to Valiant (Val84).

In fact, Ensemble Learning techniques do not provide general purpose classifiers; they only indicate the optimal way to combine multiple classifiers.

Examples of Ensemble Learning techniques include

Decision Tree
Decision Trees, being constructed from many Decision Stumps arranged in cascade, provide an initial example of Ensemble Learning;
Bagging
BootStrap AGGregatING attempts to reduce overfitting by training several classifiers on subsets of the training set and finally performing a majority vote;
Boosting
Instead of selecting purely random subsets of the training set, some of the samples that remain incorrectly classified are used;
AdaBoost
ADAptive BOOSTing (Section 5.8.2) is the best-known Ensemble Learning algorithm and the forerunner of the large family of AnyBoost classifiers;
Random ForestTM
is a BootStrap Aggregating (bagging) of Decision Trees, an Ensemble Classifier composed of several decision trees, each created from a subset of the training data and features to be analyzed, which vote by majority;
and many others.

Examples of weak classifiers widely used in the literature are Decision Stumps (AL92) associated with Haar features (Section 7.1). The Decision Stump is a binary classifier of the form

\begin{displaymath}
h(\mathbf{x}) = \left\{ \begin{array}{ll}
+1 & \quad \text...
...theta \\
-1 & \quad \text{otherwise} \\
\end{array}\right.
\end{displaymath} (5.72)

where $f(\mathbf{x})$ is a function that extracts a scalar from the sample being classified, $p=\{ +1, -1 \}$ is a parity indicating the direction of the inequality, and $\theta $ is the decision threshold (Figure 5.4).



Subsections
Paolo medici
2026-10-06