Logistic Regression

Figure 4.2: Logistic Function
Image fig_logit

There is a family of linear models that relate the dependent variable to the explanatory variables through a nonlinear function, known as generalized linear models (generalized linear model). Logistic regression belongs to this class of models when the variable $y$ is dichotomous, that is, when it can take only the values $0$ or $1$. By its nature, this type of problem is particularly important in classification problems.

For binary problems, it is possible to define the probabilities of success and failure:

\begin{displaymath}
\begin{array}{l}
P[Y=1\vert\mathbf{x}]=p(\mathbf{x}) \\
P[Y=0\vert\mathbf{x}]=1-p(\mathbf{x}) \\
\end{array}\end{displaymath} (4.118)

The response of a linear predictor of the form

\begin{displaymath}
y' = \boldsymbol\beta \cdot \mathbf{x} + \varepsilon
\end{displaymath} (4.119)

is not bounded between $0$ and $1$ and is therefore unsuitable for this purpose. It is necessary to associate the response of the linear predictor with the response of a function $g$, a function of the probability $p(\mathbf{x})$:
\begin{displaymath}
g(p(\mathbf{x}) ) = \boldsymbol\beta \cdot \mathbf{x} + b
\end{displaymath} (4.120)

where $g(p)$, the mean function, is a nonlinear function defined over $[0,1]$. $g(p)$ must be invertible, and its inverse $g^{-1}(y')$ is the link function.

A widely used model for the function $g(p)$ is the logit function, defined as:

\begin{displaymath}
logit(p) = \log \frac{p}{1-p} = \boldsymbol\beta \cdot \mathbf{x}
\end{displaymath} (4.121)

The function $\frac{p}{1-p}$, since it represents how many times more likely success is than failure, is called the odds ratio; consequently, function (4.121) represents the logarithm of the probability that an event occurs relative to the probability that the same event does not occur (log-odds).

Its inverse exists and is given by

\begin{displaymath}
\E[Y\vert\mathbf{x}] = p( \mathbf{x} ) = \frac{ e^{\boldsymb...
...dot \mathbf{x}} }{ 1 + e^{\boldsymbol\beta \cdot \mathbf{x}} }
\end{displaymath} (4.122)

and is the logistic function.

In this case, the maximum-likelihood method does not coincide with the least-squares method but with

\begin{displaymath}
\mathcal{L}(\boldsymbol\beta) = \prod_{i=1}^{n} f(y_i\vert \...
...y_i} ( \mathbf{x}_i) \left(1 - p^{y_i} ( \mathbf{x}_i) \right)
\end{displaymath} (4.123)

which gives the log-likelihood function
\begin{displaymath}
\log \mathcal{L}(\boldsymbol\beta) = \sum_{i=1}^{n} y_i (\bo...
...log \left( 1 + e^{\boldsymbol\beta \cdot \mathbf{x}_i} \right)
\end{displaymath} (4.124)

whose maximization through iterative techniques allows the parameters $\boldsymbol\beta$ to be estimated.

Paolo medici
2026-10-06