Performance Evaluation

Given a classifier trained on a particular training set (Training Set), it must be evaluated on another set (Validation Set or Certification Set). This comparison makes it possible to derive metrics for evaluating the classifier and comparing different classifiers. It is essential that performance metrics be computed on a set of samples not used during training (the validation set), in order to detect problems such as data overfitting, that is, a failure to generalize.

Once the classifier parameters have been fixed, the contingency table (Confusion Matrix) can be constructed:


    True Value
    p n
Classification p' VP FP
n' FN VN


False Positives (FP) are also referred to as false alarms. False Negatives (FN) are referred to as misses.

The table is normally used to derive performance metrics such as:

Each classifier has one or more parameters that, when modified, change the ratio between correct detections and the number of false positives. It is therefore difficult to compare two classifiers objectively because, at the same threshold, one may produce more correct detections than the other but also a larger number of false positives. To compare the performance of different binary classifiers obtained from different training sessions, curves are therefore normally used to show performance as a function of the classifier's internal threshold.

The performance curves that can be used are:

Finally, it should be noted that these metrics apply to any class of problems involving the concept of a correct or incorrect result. They are therefore applicable not only to classifiers but also, for example, to feature-point matching and other tasks.

More recently, to facilitate comparisons of classifier performance, functions have been proposed that, when applied to ROC curves, yield a single scalar representing a measure of classification quality. These functions are generally averages of samples from the ROC curve restricted to the regions of practical interest.

Paolo medici
2026-10-01