paradox #23
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Accuracy paradox

You will learn: Compare classification metrics and expected loss using explicit error costs.

Start with: Conditioning and independence

An alert system can be 99% accurate without catching a single positive case. If one percent of observations are positive, always predicting negative gets the other 99 percent right. The number is correct; its relevance depends on the system’s purpose.

The accuracy paradox arises when a high fraction of correct predictions is treated as sufficient evidence of useful classification. Keep population composition, conditional error rates, and costs visible. This experiment specifies those quantities directly rather than fitting a classifier or estimating performance from a sample.

Compare rules on the same population

Four exhaustive classification outcomes00.20.40.60.81TPFNFPTNoutcomefraction of all cases

Filled bars are correct predictions; open bars are errors. Counts below are exact expected frequencies per 10,000, so fractions of a case are allowed.

Actual classPredicted positivePredicted negative
PositiveTP 80.00000FN 20.00000
NegativeFP 495.0000TN 9405.000
MeasureClassifierAlways negative
Accuracy0.94850000.9900000
Precision0.1391304Undefined
Recall0.80000000
Balanced accuracy0.87500000.5
F10.23703700
Expected loss per case0.14950000.5000000

A false alarm costs one unit, a missed positive costs 50, and correct decisions cost zero. The lower-loss rule is the classifier. Costs are choices, not estimated facts. Precision requires predicted positives; recall requires actual positives; balanced accuracy here requires both classes.

A synthetic population with specified conditional rates. This panel does not fit a classifier, infer its calibration, or claim that every pair of rates is attainable by changing one threshold.

The default has one percent positives, recall 80%, and false-positive rate 5%. Per 10,000 expected observations, 100 are actual positives: 80 true positives and 20 false negatives. Among the other 9,900, there are 495 false positives and 9,405 true negatives.

Accuracy is (80 + 9,405)/10,000 = 94.85%. Always predicting negative achieves 99%, so accuracy prefers the rule missing all 100 positives. This is not an arithmetic error: accuracy assigns the same unit cost to each false positive and false negative, and the always-negative rule makes fewer total errors here.

Different metrics answer different questions

Recall is TP/(TP + FN): what fraction of actual positives are found? Precision is TP/(TP + FP): what fraction of predicted positives are correct? Default precision is 80/575, about 13.91%. High recall does not guarantee high precision because false positives come from a much larger negative population.

Specificity is TN/(TN + FP), or 95%. Balanced accuracy averages recall and specificity, giving 87.5%. It weights actual classes equally regardless of prevalence. That can be useful, but remains a choice of objective.

F1 is 2TP/(2TP + FP + FN). It combines precision and recall while excluding true negatives, so it does not describe all four confusion-matrix cells. No metric removes the need to specify relevant errors.

Put error costs into the comparison

Let a false positive cost one unit and a missed positive cost c units. Expected loss per observation is c times false-negative probability plus false-positive probability. At default c = 50, the model’s loss is 50 × 0.002 + 0.0495 = 0.1495 units. Always predicting negative costs 50 × 0.01 = 0.5 units.

Under these costs, the alert model has lower expected loss despite lower accuracy. At c = 1, loss is one minus accuracy, and always-negative wins. For the default rates, alerts have lower loss exactly when c > 0.0495/0.008 = 6.1875. This compares two rules; it does not prove either is optimal among all classifiers.

Make a prediction

If prevalence rises while recall and false-positive rate stay fixed, must precision stay unchanged?

Explore the answer

No. True-positive mass grows and false-positive mass shrinks. At prevalence 10%, recall 80%, and false-positive rate 5%, precision is 0.08/(0.08 + 0.045) = 64%. Conditional rates and the composition of predicted positives are different quantities.

Boundaries and practical limits

Without predicted positives, precision is undefined. Recall is undefined without actual positives; specificity is undefined without actual negatives. Balanced accuracy here requires both classes. Fractional expected counts describe probabilities scaled to a reference population, not a realized dataset.

Changing a score threshold often changes recall and false-positive rate together. Independent sliders do not claim every pair is attainable by one fitted model. Costs can vary, deployment populations can differ, and estimated rates have uncertainty. The display isolates specified rates and costs without resolving those additional questions.

Reference and connection

Google’s Machine Learning Crash Course: accuracy, recall, precision defines the metrics. The counts and cost comparison follow the model here. Continue to base rate neglect to examine positive predictive value for a rare condition.

Reset all settings