The Boosting problem can be generalized and viewed as a problem in which predictors
must be found that minimize the global cost function:
| (5.92) |
From an analytical perspective, AdaBoost is an example of a coordinate-wise gradient-descent optimizer that minimizes the potential function
, optimizing one coefficient
at a time (LS10), as shown by Equation (5.81).
The following is a non-exhaustive list of AdaBoost variants that highlights some of the distinctive features of this technique:
AdaBoost can also be extended to classifiers with abstention, for which the possible outputs are
.
Extending Definition (5.84), for simplicity let
denote failures,
abstentions, and
successes of classifier
.
In this case as well, attains its minimum with the same value
as in the case without abstention; see (5.90),
and with this choice
would be
| (5.93) |
However, Freund and Shapire proposed a more conservative choice of :
| (5.94) |
Real AdaBoost generalizes the preceding case and, above all, generalizes the same extended additive model (FHT00).
Instead of using dichotomous hypotheses and assigning them a weight
, it directly seeks the feature
that minimizes Equation (5.81).
Real AdaBoost allows weak classifiers to provide the probability distribution
, namely the probability that class
is actually
given the observation of feature
.
Given a probability distribution , the feature
that minimizes Equation (5.81) is
Real AdaBoost can also be used with a discrete classifier such as the Decision Stump.
By directly applying Equation (5.95) to the two possible output states of the Decision Stump (the minimum of can nevertheless be obtained easily algebraically), the classifier responses must take the values
| (5.97) |
| (5.98) |
Gentle AdaBoost further generalizes the concept of Ensemble Learning to an additive model (FHT00) by using a regression with steps typical of Newton methods:
| (5.99) |
The hypothesis to be added to the additive model at iteration
is selected from all possible hypotheses
as the one that optimizes a weighted least-squares regression:
| (5.101) |
For historical reasons, AdaBoost does not explicitly exhibit a statistical formalism.
The first observation is that the response of the AdaBoost classifier is not a probability, since it is not bounded between .
In addition to this problem, which is partially resolved by Real AdaBoost, minimizing the loss function (5.80) does not appear to be a statistical approach, unlike maximizing the likelihood.
Nevertheless, it can be shown that the AdaBoost cost function maximizes a function very similar to the Bernoulli log-likelihood.
For these reasons, AdaBoost can be extended to logistic regression theory, described in Section 5.4.
Additive logistic regression takes the form
The problem is to find a suitable loss function for this representation, that is, to identify an AdaBoost variant that exactly maximizes the Bernoulli log-likelihood (FHT00).
Maximizing the likelihood of (5.103) is equivalent to minimizing the log-loss
| (5.104) |
LogitBoost was the first method to extend AdaBoost to the logistic optimization problem for a function under the cost function
, maximizing the Bernoulli log-likelihood using Newton-type iterations.
The weights associated with each sample follow directly from the probability distribution
| (5.105) |
This behavior can be represented by a cost function of the form
| (5.106) |
Paolo medici