Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Empirical CDF
You will learn: Retain ties and compute cumulative probabilities and inverse-CDF quantiles from a sample.
Start with: Reading a distribution
What fraction of the observations are at or below a threshold? The empirical cumulative distribution function answers that question directly. It puts equal probability on each observed row, retains repeated values, and requires no histogram bins or fitted density. Start with reading a distribution if cumulative probability and quantiles are new.
For observed values x₁,…,xₙ, define:
The indicator contributes one when the observation meets the threshold. Dividing by n turns its count into a fraction. A finite dataset defines this distribution regardless of how it was collected. Treating it as evidence about a population is a separate step.
Edit up to 50 numbers between −1000 and 1000. The default is a synthetic five-value sample. These values are included in the copyable scenario URL.
Filled circles include the jump; open circles mark the left limit. Each observation contributes 1/5. Gold marks the inverse-CDF quantile at p=0.50: 2.
| At threshold 2 | Count | Fraction |
|---|---|---|
| Strictly below | 1 / 5 | 0.200000 |
| Exactly equal | 2 / 5 | 0.400000 |
| At or below | 3 / 5 | 0.600000 |
Change bins without changing the empirical distribution
Histogram heights are bin fractions, not densities. The ECDF retains exact observed positions and ties without choosing bins. All observations remain in both calculations.
Inspect every distinct value and its probability
| Value | Multiplicity | Mass | Left limit | Fₙ(value) |
|---|---|---|---|---|
| 1 | 1 | 0.2000 | 0.0000 | 0.2000 |
| 2 | 2 | 0.4000 | 0.2000 | 0.6000 |
| 4 | 1 | 0.2000 | 0.6000 | 0.8000 |
| 7 | 1 | 0.2000 | 0.8000 | 1.0000 |
Five observations, four jumps
The default sample is 1, 2, 2, 4, 7. Its empirical distribution puts mass 1/5 at 1, 2/5 at 2, 1/5 at 4, and 1/5 at 7. Keeping only distinct values and giving each mass 1/4 would change the question by discarding multiplicities.
Just below 2, only the value 1 has been counted, so the CDF’s left limit is 1/5. At exactly 2, the two tied observations join it: Fₙ(2)=3/5. Thus Pₙ(X<2)=0.2, Pₙ(X=2)=0.4, and Pₙ(X≤2)=0.6. The filled circle includes the jump; the open circle marks the excluded left limit. Between observations the curve stays flat, and beyond the maximum it stays at one.
Make a prediction
If you add one more observation equal to 2, does its empirical probability become 3/5?
Explore the answer
No. The numerator and denominator both change. The new sample has six observations and three twos, so the mass at 2 is 3/6. Its CDF value is 4/6 because the observation equal to 1 also counts. Edit the sample to verify both quantities.
Invert the staircase carefully
For 0<p≤1, this lesson uses the generalized inverse:
where the parenthesized index denotes the sorted observations, including repeats. In the default sample, Qₙ(0.5)=2 and Qₙ(0.6)=2, but Qₙ(0.61)=4. At a jump, many probability levels share a quantile. It is generally false that Fₙ(Qₙ(p))=p; here Fₙ(Qₙ(0.5))=0.6.
Different software can use interpolated sample quantiles. For the two-value sample 1,9, this generalized-inverse median is 1, while averaging the middle pair gives 5. Both are conventions used for sample summaries; use the convention appropriate to the calculation and state it. The control excludes p=0 because the displayed infimum definition would then range over every real t, rather than return the observed minimum.
Make a prediction
If all observations equal 3, where is the empirical median and how large is the CDF jump?
Explore the answer
The median is 3 and the jump there is one. Repeated observations contribute separately even when their displayed positions coincide. There is no positive spread to smooth into a bell.
Change the picture without changing the sample
Histogram bins group nearby values and can alter the appearance of a distribution. The histogram here uses fractions per bin, not density heights. Its bin fractions sum to one. Increasing its bin count neither moves an observation nor changes the ECDF or its inverse quantiles.
For a fixed threshold t, the empirical fraction is also the sample average of event indicators. Under independent identically distributed sampling from a population CDF F, those indicators have mean F(t), so the Law of Large Numbers explains convergence at that fixed threshold. This argument needs a sampling model; a convenient sample, repeated measurements from one individual, or dependent time gaps does not become representative merely because its ECDF is easy to draw.
The empirical distribution assigns zero mass outside observed values. That does not establish that the population cannot produce something new or more extreme. The bootstrap samples from this finite distribution and inherits its limitations. In the earthquake-arrival case study, a gap ECDF summarizes the complete observed gaps while the prose keeps population and dependence assumptions separate.
Sources
Random Services: samples and statistics develops the empirical distribution. Its order-statistics chapter gives the inverse-CDF sample-quantile convention. The five-observation example here exposes every contributing row for direct verification.
Return to learning paths and foundations.