Soft Margin SVM

In real-world applications, a margin does not always exist; that is, the classes are not always linearly separable in feature space by a hyperplane. The concept underlying the Soft Margin overcomes this limitation by introducing an additional variable $\xi$ for each sample, thereby relaxing (slack) the margin constraint
\begin{displaymath}
\begin{array}{l}
y_i (\mathbf{w} \cdot \mathbf{x}_i + b) \ge 1 - \xi_i \\
\xi_i \ge 0, \forall i
\end{array}\end{displaymath} (5.37)

The parameter $\xi$ represents the slackness associated with the sample. When $0<\xi\le 1$, the sample is correctly classified but lies within the margin region. When $\xi>1$, the sample enters the decision region of the opposite class and is therefore misclassified.

It follows immediately from the definition that

\begin{displaymath}
\xi_i \ge \max \left( 0, 1-y_if(\mathbf{x}_i) \right)
\end{displaymath} (5.38)

At the optimum, the slack variables equal the minimum feasible value, and therefore
\begin{displaymath}
\xi_i= \max \left( 0, 1-y_if(\mathbf{x}_i) \right).
\end{displaymath} (5.39)

To find a separation hyperplane that is optimal in some sense, the cost function to be minimized must also account for the distance between the sample and the margin:

\begin{displaymath}
\min \frac{1}{2} \Vert \mathbf{w} \Vert^2 + C \sum \xi_i
\end{displaymath} (5.40)

subject to constraints (5.37). The parameter $C$ is a degree of freedom of the problem that indicates how much a sample must pay for violating the margin constraint. When $C$ is small, the margin is wide, whereas when $C$ approaches infinity, the formulation reduces to the Hard Margin SVM formulation discussed above.

Each sample $\mathbf{x}_i$ can fall into one of three possible states:

The Lagrangian of system (5.40), with the constraints introduced by variables $\xi$, is

\begin{displaymath}
\mathcal{L}(\mathbf{w},b,\xi,\alpha) = \frac{1}{2} \Vert\mat...
...} \cdot \mathbf{x}_i + b) - 1 + \xi_i) - \sum_i \gamma_i \xi_i
\end{displaymath} (5.41)

With the increased number of constraints, the dual variables are both $\bm{\alpha}$ and $\bm{\gamma}$.

The remarkable result is that, after taking the derivatives, the dual formulation of (5.41) becomes exactly the same as the dual formulation of the Hard Margin case: the variables $\xi_i$ do not appear in the dual formulation, and the only difference between the Hard Margin and Soft Margin cases lies in the constraint on the parameters $\alpha_i$, which in this case are bounded by

\begin{displaymath}
0 \le \alpha_i \le C
\end{displaymath} (5.42)

rather than by the simple inequality $\alpha_i \ge 0$. The great advantage of this formulation is precisely the simplicity of the constraints and the fact that it reduces the Hard Margin case to a special case ($C=\infty$) of the Soft Margin. The constant $C$ is an upper bound on the value that the $\alpha_i$ can take.

Paolo medici
2026-10-01