distribution #28
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Empirical CDF

You will learn: Retain ties and compute cumulative probabilities and inverse-CDF quantiles from a sample.

Start with: Reading a distribution

What fraction of the observations are at or below a threshold? The empirical cumulative distribution function answers that question directly. It puts equal probability on each observed row, retains repeated values, and requires no histogram bins or fitted density. Start with reading a distribution if cumulative probability and quantiles are new.

For observed values x₁,…,xₙ, define:

Fn(t)=1n∑i=1n1{xi≤t}.F_n(t)=\frac{1}{n}\sum_{i=1}^n \mathbf{1}\{x_i\le t\}.

The indicator contributes one when the observation meets the threshold. Dividing by n turns its count into a fraction. A finite dataset defines this distribution regardless of how it was collected. Treating it as evidence about a population is a separate step.

Edit up to 50 numbers between −1000 and 1000. The default is a synthetic five-value sample. These values are included in the copyable scenario URL.

Right-continuous empirical CDF with visible jumps at ties00.20.40.60.81012345678observed valueempirical cumulative probability

Filled circles include the jump; open circles mark the left limit. Each observation contributes 1/5. Gold marks the inverse-CDF quantile at p=0.50: 2.

At threshold 2CountFraction
Strictly below1 / 50.200000
Exactly equal2 / 50.400000
At or below3 / 50.600000

Change bins without changing the empirical distribution

Histogram bin fractions from the same complete sample00.10.20.30.4012345678observed valuefraction in bin

Histogram heights are bin fractions, not densities. The ECDF retains exact observed positions and ties without choosing bins. All observations remain in both calculations.

Inspect every distinct value and its probability
ValueMultiplicityMassLeft limitFₙ(value)
110.20000.00000.2000
220.40000.20000.6000
410.20000.60000.8000
710.20000.80001.0000
A finite dataset defines a distribution without assuming independent sampling. Inferring an underlying population distribution requires assumptions about how the data were collected.

Five observations, four jumps

The default sample is 1, 2, 2, 4, 7. Its empirical distribution puts mass 1/5 at 1, 2/5 at 2, 1/5 at 4, and 1/5 at 7. Keeping only distinct values and giving each mass 1/4 would change the question by discarding multiplicities.

Just below 2, only the value 1 has been counted, so the CDF’s left limit is 1/5. At exactly 2, the two tied observations join it: Fₙ(2)=3/5. Thus Pₙ(X<2)=0.2, Pₙ(X=2)=0.4, and Pₙ(X≤2)=0.6. The filled circle includes the jump; the open circle marks the excluded left limit. Between observations the curve stays flat, and beyond the maximum it stays at one.

Make a prediction

If you add one more observation equal to 2, does its empirical probability become 3/5?

Explore the answer

No. The numerator and denominator both change. The new sample has six observations and three twos, so the mass at 2 is 3/6. Its CDF value is 4/6 because the observation equal to 1 also counts. Edit the sample to verify both quantities.

Invert the staircase carefully

For 0<p≤1, this lesson uses the generalized inverse:

Qn(p)=inf⁡{t:Fn(t)≥p}=x(⌈np⌉),Q_n(p)=\inf\{t:F_n(t)\ge p\}=x_{(\lceil np\rceil)},

where the parenthesized index denotes the sorted observations, including repeats. In the default sample, Qₙ(0.5)=2 and Qₙ(0.6)=2, but Qₙ(0.61)=4. At a jump, many probability levels share a quantile. It is generally false that Fₙ(Qₙ(p))=p; here Fₙ(Qₙ(0.5))=0.6.

Different software can use interpolated sample quantiles. For the two-value sample 1,9, this generalized-inverse median is 1, while averaging the middle pair gives 5. Both are conventions used for sample summaries; use the convention appropriate to the calculation and state it. The control excludes p=0 because the displayed infimum definition would then range over every real t, rather than return the observed minimum.

Make a prediction

If all observations equal 3, where is the empirical median and how large is the CDF jump?

Explore the answer

The median is 3 and the jump there is one. Repeated observations contribute separately even when their displayed positions coincide. There is no positive spread to smooth into a bell.

Change the picture without changing the sample

Histogram bins group nearby values and can alter the appearance of a distribution. The histogram here uses fractions per bin, not density heights. Its bin fractions sum to one. Increasing its bin count neither moves an observation nor changes the ECDF or its inverse quantiles.

For a fixed threshold t, the empirical fraction is also the sample average of event indicators. Under independent identically distributed sampling from a population CDF F, those indicators have mean F(t), so the Law of Large Numbers explains convergence at that fixed threshold. This argument needs a sampling model; a convenient sample, repeated measurements from one individual, or dependent time gaps does not become representative merely because its ECDF is easy to draw.

The empirical distribution assigns zero mass outside observed values. That does not establish that the population cannot produce something new or more extreme. The bootstrap samples from this finite distribution and inherits its limitations. In the earthquake-arrival case study, a gap ECDF summarizes the complete observed gaps while the prose keeps population and dependence assumptions separate.

Sources

Random Services: samples and statistics develops the empirical distribution. Its order-statistics chapter gives the inverse-CDF sample-quantile convention. The five-observation example here exposes every contributing row for direct verification.

Return to learning paths and foundations.

Reset all settings