Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Accuracy paradox
You will learn: Compare classification metrics and expected loss using explicit error costs.
Start with: Conditioning and independence
An alert system can be 99% accurate without catching a single positive case. If one percent of observations are positive, always predicting negative gets the other 99 percent right. The number is correct; its relevance depends on the system’s purpose.
The accuracy paradox arises when a high fraction of correct predictions is treated as sufficient evidence of useful classification. Keep population composition, conditional error rates, and costs visible. This experiment specifies those quantities directly rather than fitting a classifier or estimating performance from a sample.
Compare rules on the same population
Filled bars are correct predictions; open bars are errors. Counts below are exact expected frequencies per 10,000, so fractions of a case are allowed.
| Actual class | Predicted positive | Predicted negative |
|---|---|---|
| Positive | TP 80.00000 | FN 20.00000 |
| Negative | FP 495.0000 | TN 9405.000 |
| Measure | Classifier | Always negative |
|---|---|---|
| Accuracy | 0.9485000 | 0.9900000 |
| Precision | 0.1391304 | Undefined |
| Recall | 0.8000000 | 0 |
| Balanced accuracy | 0.8750000 | 0.5 |
| F1 | 0.2370370 | 0 |
| Expected loss per case | 0.1495000 | 0.5000000 |
A false alarm costs one unit, a missed positive costs 50, and correct decisions cost zero. The lower-loss rule is the classifier. Costs are choices, not estimated facts. Precision requires predicted positives; recall requires actual positives; balanced accuracy here requires both classes.
The default has one percent positives, recall 80%, and false-positive rate 5%. Per 10,000 expected observations, 100 are actual positives: 80 true positives and 20 false negatives. Among the other 9,900, there are 495 false positives and 9,405 true negatives.
Accuracy is (80 + 9,405)/10,000 = 94.85%. Always predicting negative achieves 99%, so accuracy prefers the rule missing all 100 positives. This is not an arithmetic error: accuracy assigns the same unit cost to each false positive and false negative, and the always-negative rule makes fewer total errors here.
Different metrics answer different questions
Recall is TP/(TP + FN): what fraction of actual positives are found? Precision is TP/(TP + FP): what fraction of predicted positives are correct? Default precision is 80/575, about 13.91%. High recall does not guarantee high precision because false positives come from a much larger negative population.
Specificity is TN/(TN + FP), or 95%. Balanced accuracy averages recall and specificity, giving 87.5%. It weights actual classes equally regardless of prevalence. That can be useful, but remains a choice of objective.
F1 is 2TP/(2TP + FP + FN). It combines precision and recall while excluding true negatives, so it does not describe all four confusion-matrix cells. No metric removes the need to specify relevant errors.
Put error costs into the comparison
Let a false positive cost one unit and a missed positive cost c units. Expected loss per observation is c times false-negative probability plus false-positive probability. At default c = 50, the model’s loss is 50 × 0.002 + 0.0495 = 0.1495 units. Always predicting negative costs 50 × 0.01 = 0.5 units.
Under these costs, the alert model has lower expected loss despite lower accuracy. At c = 1, loss is one minus accuracy, and always-negative wins. For the default rates, alerts have lower loss exactly when c > 0.0495/0.008 = 6.1875. This compares two rules; it does not prove either is optimal among all classifiers.
Make a prediction
If prevalence rises while recall and false-positive rate stay fixed, must precision stay unchanged?
Explore the answer
No. True-positive mass grows and false-positive mass shrinks. At prevalence 10%, recall 80%, and false-positive rate 5%, precision is 0.08/(0.08 + 0.045) = 64%. Conditional rates and the composition of predicted positives are different quantities.
Boundaries and practical limits
Without predicted positives, precision is undefined. Recall is undefined without actual positives; specificity is undefined without actual negatives. Balanced accuracy here requires both classes. Fractional expected counts describe probabilities scaled to a reference population, not a realized dataset.
Changing a score threshold often changes recall and false-positive rate together. Independent sliders do not claim every pair is attainable by one fitted model. Costs can vary, deployment populations can differ, and estimated rates have uncertainty. The display isolates specified rates and costs without resolving those additional questions.
Reference and connection
Google’s Machine Learning Crash Course: accuracy, recall, precision defines the metrics. The counts and cost comparison follow the model here. Continue to base rate neglect to examine positive predictive value for a rare condition.