paradox #12
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Inspection paradox

You will learn: Identify how time weighting changes selected interval lengths and remaining waits.

Start with: Expectation and variance · Conditioning and independence

A randomly chosen interval and the interval containing a randomly chosen time are different observations. Longer intervals occupy more of the timeline and are more likely to contain a uniform-time inspection. Review weighted expectation and conditioning.

This experiment uses a finite repeating schedule so the sampling rule and denominator can be inspected exactly. It needs no claim that real buses have independent, identical gaps or that an infinite-time limit has already been reached.

This synthetic schedule repeats 1 interval of length 2, then 1 of length 18. Time units are arbitrary. Arrivals occur at interval boundaries; an observer arrives independently of the schedule.

A repeating arrival schedule with the inspected time and next arrival05101520time in one cycle

Outlined blocks: type A. Filled blue blocks: type B. Width is duration. The solid red marker is the inspection; the dashed black line is the next arrival. All intervals in one cycle are included.

One sampled inspectionValue
Selected interval2, type B
Inspection time19.64039
Interval start / next arrival2.000000 / 20.00000
Remaining wait0.3596107
Exact comparisonUniform timeUniform interval, then time
Probability selected interval is A0.10000000.5000000
Mean selected interval length16.4000010.00000
Mean remaining wait8.2000005.000000
P(wait > 10)0.40000000.2222222

Mean interval length across arrivals is 10.00000. Uniform-time sampling weights an interval by its duration; uniform-interval sampling gives every interval equal weight before drawing its interior time. The selected protocol changes the example marker, while both exact calculations remain visible.

An exact finite repeating-schedule experiment, not a fitted bus timetable or a claim about a specific transit system. The two sampling protocols have different denominators.

Two protocols, two answers

The default cycle contains one interval of length two and one of length eighteen, for total duration twenty. An interval chosen uniformly from the two has mean length ten. A uniformly chosen time falls in the long interval with probability 18/20=0.9, so the mean length containing that time is 0.1·2+0.9·18=16.4.

Within a selected interval of length L, a uniform interior time has mean remaining wait L/2. Thus the uniform-time observer waits 8.2 on average, while choosing an interval uniformly and then an interior time gives 5. Neither average is an arithmetic mistake; they answer different sampling questions.

At threshold ten, a uniform-time observer waits longer than ten only in the first eight time units of the eighteen-unit interval. The probability is 8/20=0.4. Under uniform-interval sampling it is (1/2)·(8/18)=2/9≈0.222222.

Make a prediction

Is the average remaining wait always half the average interval length across arrivals?

Explore the answer

It is under the uniform-interval-then-uniform-interior protocol. Uniform-time inspection instead favors long intervals before taking their average half-length, so the original unweighted interval average is the wrong denominator.

Length weighting in a formula

Let pⱼ be the fraction of intervals with length Lⱼ, where every length is positive. For uniform time within a repeating cycle, the probability of selecting type j is pⱼLⱼ/E[L]. Therefore:

E[Linspected]=E[L2]E[L],E[R]=E[L2]2E[L]=E[L]2+Var⁡(L)2E[L].E[L_{\mathrm{inspected}}]=\frac{E[L^2]}{E[L]},\qquad E[R]=\frac{E[L^2]}{2E[L]}=\frac{E[L]}2+\frac{\operatorname{Var}(L)}{2E[L]}.

The extra term measures the effect of variation in interval lengths. Equal lengths make it zero. Increasing variation while holding the mean fixed can raise the average remaining wait substantially.

For an extreme finite example, use nine intervals of length one and one interval of length ninety-one. The mean interval is still ten, but the uniform-time remaining wait becomes (9·1²+91²)/(2·100)=41.45. It can exceed the mean interval across arrivals. No observer waits longer than the particular interval they are in; the apparent contradiction comes from comparing differently weighted populations.

The survival probability is equally direct. In an interval of length L, the portion of time with more than x remaining is max(L−x,0). Adding those portions and dividing by total cycle duration gives:

P(R>x)=E[(L−x)+]E[L],x≥0.P(R>x)=\frac{E[(L-x)_+]}{E[L]},\qquad x\ge0.

These finite-cycle formulas follow from lengths on the timeline, independent of the order in which the interval types appear within the repeated cycle.

A schedule is part of the model

For a constant ten-unit schedule, uniform-time remaining wait is five. If arrivals instead follow a homogeneous Poisson process, exponential gaps have second moment twice the squared mean, and the equilibrium remaining wait equals the mean gap. These are different models, even when both average ten units between arrivals.

For a general renewal model, stationary or suitable long-run sampling assumptions are needed to use the equilibrium length-weighted formulas. The observer must also be sampled independently of arrivals. A passenger who checks a timetable or changes arrival time after a missed service has a different protocol. Real data also require care about missed events, observation boundaries, and incomplete intervals.

Make a prediction

Does this paradox show that sampling is dishonest or that the reported mean is false?

Explore the answer

No. Time sampling and interval sampling can both be correctly implemented. The problem is treating their means as interchangeable without identifying what was sampled and how it was weighted.

See HH versus HT for a different waiting-time mechanism: overlapping patterns change the state remembered by a search.

Sources

Random Services’ renewal-process introduction develops the inspection-interval distinction. The finite-cycle probabilities, tail areas, and two numerical examples above are direct length calculations for the explicitly specified repeating schedule. The Poisson comparison uses the exponential model.

Reset all settings