Given a classifier trained on a particular training set (Training Set), it must be evaluated on another set (Validation Set or Certification Set). This comparison makes it possible to derive metrics for evaluating the classifier and comparing different classifiers. It is essential that performance metrics be computed on a set of samples not used during training (the validation set), in order to detect problems such as data overfitting, that is, a failure to generalize.
Once the classifier parameters have been fixed, the contingency table (Confusion Matrix) can be constructed:
| True Value | |||
| p | n | ||
| Classification | p' | TP | FP |
| n' | FN | TN | |
False Positives (FP) are also called false alarms, while False Negatives (FN) are called misses.
The following performance metrics are typically derived from the table:
Each classifier has one or more parameters that, when modified, change the ratio between correct detections and the number of false positives. It is therefore difficult to compare two classifiers objectively: at the same threshold, one may produce more correct detections than the other but also a greater number of false positives. To compare the performance of different binary classifiers obtained from different training sessions, curves are normally used to show performance as this internal classifier threshold varies.
The performance curves that may be encountered include:
Finally, it should be noted that these metrics apply to any class of problems involving the concept of a correct or incorrect result. They therefore apply not only to classifiers but also, for example, to feature-point matching and other tasks.
More recently, to enable more streamlined comparisons of classifier performance, functions have been proposed that, when applied to ROC curves, yield a single scalar representing a classification-quality score. These functions are usually averages of samples from the ROC curve over the regions of practical interest.