Logistic regression is one of the simplest and most widely used linear classifiers in statistics and machine learning. Unlike LDA, which derives from probabilistic assumptions about class distributions, logistic regression directly models the posterior probability of the categories and is therefore classified as a discriminative method.
As with other linear classifiers, a discriminant function of the form is considered
| (5.17) |
where is the weight vector and
is the bias term.
In the binary case, assuming labels
,
the probability that sample
belongs to the positive class is modeled by the logistic function
| (5.18) |
whereas the probability of the negative class is
| (5.19) |
The logistic function transforms the real value of the discriminant function into a probability between and
.
The separating surface between the two classes corresponds to the condition
| (5.20) |
that is,
| (5.21) |
which coincides with a linear hyperplane.
Training consists of estimating the parameters and
that maximize the likelihood of the observed data.
Given a training set
,
the log-likelihood is written as
| (5.22) |
where
| (5.23) |
Maximizing the log-likelihood is equivalent to minimizing the cost function
| (5.24) |
known as the cross-entropy loss or log-loss.
Unlike LDA, logistic regression makes no assumptions about the class distributions and can therefore also be applied when the observations do not follow a Gaussian distribution. On the other hand, like all linear classifiers, it can correctly separate only classes that are linearly separable in feature space.
Logistic regression can be interpreted as a probabilistic version of linear classifiers.
The value represents the logarithm of the ratio between the probabilities of the two classes
| (5.25) |
and is often called the log-odds or logit.
Paolo medici