Given a loss function , the expected risk of a classifier is defined as
| (5.54) |
Since the data distribution is generally unknown, in practice the risk is estimated from a finite set of training samples:
| (5.55) |
This quantity is called the empirical risk. The principle of Empirical Risk Minimization (ERM) therefore consists in finding the model that minimizes the empirical risk on the training data.
In practice, empirical risk minimization is often accompanied by a regularization term that limits the complexity of the model:
| (5.56) |
This formulation makes it possible to interpret the Soft Margin SVM introduced in Section 5.5 as a special case of empirical risk minimization with regularization.