In real-world applications, a margin does not always exist; that is, the classes are not always linearly separable in feature space by a hyperplane.
The concept underlying the Soft Margin overcomes this limitation by introducing an additional variable
for each sample, thereby relaxing (slackening) the margin constraint:
 |
(5.37) |
The parameter
represents the slackness associated with the sample.
When
, the sample is correctly classified but lies within the margin region.
When
, the sample enters the decision space of the opposite class and is therefore misclassified.
The definition immediately implies that
 |
(5.38) |
At the optimum, the slack variables take the minimum feasible value, and therefore
 |
(5.39) |
To find a separating hyperplane that is in some sense optimal, the cost function to be minimized must also account for the distance between the sample and the margin:
 |
(5.40) |
subject to constraints (5.37).
The parameter
is a degree of freedom that indicates how much a sample must pay for violating the margin constraint.
When
is small, the margin is wide, whereas when
approaches infinity, the formulation reduces to the Hard Margin SVM formulation introduced earlier.
Each sample
can be in one of three possible states:
- it may lie beyond the margin
and consequently not contribute to the function;
- it may lie on or inside the margin and contribute to the solution as a support vector;
- it may finally lie inside the margin and be penalized in proportion to its deviation from the hard constraints.
The Lagrangian of system (5.40), with the constraints introduced by the variables
, is
 |
(5.41) |
As the number of constraints increases, the dual variables are both
and
.
The remarkable result is that, after applying the derivatives, the dual formulation of (5.41) becomes exactly identical to the dual formulation of the Hard Margin case:
the variables
do not appear in the dual formulation,
and the only difference between the Hard Margin and Soft Margin cases lies in the constraint on the parameters
, which in this case are bounded by
 |
(5.42) |
instead of the simple inequality
.
The main advantage of this formulation is precisely the simplicity of the constraints and the fact that it reduces the Hard Margin case to a particular case (
) of the Soft Margin.
The constant
is an upper bound on the value that the
can take.
Paolo medici
2026-10-06