In real-world applications, a margin does not always exist; that is, the classes are not always linearly separable in feature space by a hyperplane.
The concept underlying the Soft Margin overcomes this limitation by introducing an additional variable
for each sample, thereby relaxing (slack) the margin constraint
 |
(5.37) |
The parameter
represents the slackness associated with the sample.
When
, the sample is correctly classified but lies within the margin region.
When
, the sample enters the decision region of the opposite class and is therefore misclassified.
It follows immediately from the definition that
 |
(5.38) |
At the optimum, the slack variables equal the minimum feasible value, and therefore
 |
(5.39) |
To find a separation hyperplane that is optimal in some sense, the cost function to be minimized must also account for the distance between the sample and the margin:
 |
(5.40) |
subject to constraints (5.37).
The parameter
is a degree of freedom of the problem that indicates how much a sample must pay for violating the margin constraint.
When
is small, the margin is wide, whereas when
approaches infinity, the formulation reduces to the Hard Margin SVM formulation discussed above.
Each sample
can fall into one of three possible states:
- it may lie beyond the margin
and consequently not contribute to the function;
- it may lie on the margin or within the margin and contribute to the solution as a support vector;
- it may finally lie within the margin and be penalized in proportion to its deviation from the hard constraints.
The Lagrangian of system (5.40), with the constraints introduced by variables
, is
 |
(5.41) |
With the increased number of constraints, the dual variables are both
and
.
The remarkable result is that, after taking the derivatives, the dual formulation of (5.41) becomes exactly the same as the dual formulation of the Hard Margin case:
the variables
do not appear in the dual formulation,
and the only difference between the Hard Margin and Soft Margin cases lies in the constraint on the parameters
, which in this case are bounded by
 |
(5.42) |
rather than by the simple inequality
.
The great advantage of this formulation is precisely the simplicity of the constraints and the fact that it reduces the Hard Margin case to a special case (
) of the Soft Margin.
The constant
is an upper bound on the value that the
can take.
Paolo medici
2026-10-01