In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Reading a distribution

A random variable assigns a number to an outcome. A distribution tells us how probability is allocated across those numbers. Before interpreting a chart, identify whether the variable is discrete or continuous, and read the vertical axis.

A die has discrete outcomes. Its probability mass function, or PMF, gives the probability of each face. Those six probabilities add to one. A normal model is continuous: its density curve allocates probability through area, not height.

Area under a standard normal density00.10.20.30.4−4−3−2−101234zdensity

P(-1 ≤ Z ≤ 1) = 0.682689 (68.269%).

CDF at upper bound: 0.841345. Upper tail: 0.158655.

Shading represents probability; curve height represents density. Equal bounds give zero area, even where the curve is high.

Height is not probability

The chart shows a standard normal density: mean zero and standard deviation one. The height near zero is about 0.399. That does not mean a 39.9% chance of exactly zero. Every individual point has probability zero in this continuous model; an interval can have positive probability.

Set the bounds to −1 and 1. The area between them is about 0.6827. Change them to −2 and 2 and it becomes about 0.9545. Equal lower and upper bounds give zero area.

A density can exceed one. A uniform distribution over an interval of width 0.1 has density 10 throughout that interval, but its total area is still 10 × 0.1 = 1.

Accumulate from the left

The cumulative distribution function, or CDF, is F(x) = P(X ≤ x). It starts near zero and ends near one. For a continuous variable, the probability between a and b is F(b) − F(a).

For an integer count, endpoints require care: P(a ≤ X ≤ b) = F(b) − F(a−1). Subtracting F(a) would accidentally remove the probability at a. The Poisson lesson lets you compare point and cumulative probabilities.

Make a prediction

If F(1.645) is about 0.95 for a standard normal variable, does the interval from −1.645 to 1.645 contain 95%?

Explore the answer

No. Symmetry puts about 5% below −1.645 and 5% above 1.645, leaving about 90% in the middle. A central 95% interval uses approximately ±1.96.

Read in the opposite direction

A quantile asks for a value corresponding to a cumulative probability. The 0.95 quantile of a standard normal distribution is about 1.645. For a discrete distribution, the CDF jumps, so a quantile need not have cumulative probability exactly equal to the requested level: use the smallest value reaching or exceeding it.

Move from probability to a quantile

Read the CDF in the opposite direction00.20.40.60.81−4−3−2−101234valuecumulative probability

Requested p = 0.950. Quantile q = 1.644854. P(X<q) = 0.950000; P(X≤q) = 0.950000.

Dashed horizontal: requested probability. Vertical: the smallest value whose CDF reaches that probability. Filled circles include their endpoints; open circles exclude them. The normal curve is continuous. Discrete and empirical quantiles use inverse-CDF steps, without interpolation.

Choose the fair die and request p=0.95. The inverse CDF returns 6, where F(6)=1, because F(5)=5/6 is still below 0.95. The quantile reaches or exceeds the requested probability; it need not equal it.

Now choose the nine observed values and request p=0.25. The inverse empirical CDF returns −0.4: one observation lies strictly below it and three lie at or below it. This convention uses an observed value rather than interpolating between neighboring observations. Software can use other sample-quantile conventions, so state which definition you need.

Make a prediction

Can an empirical CDF jump by more than 1/n at one value?

Explore the answer

Yes. Tied observations share a location, so their probability masses add. Here −0.4 occurs twice and produces a jump of 2/9. Every observation still contributes exactly 1/9.

Separate model and sample

A theoretical curve describes a model. A histogram describes observed or simulated values and changes when you resample. Changing sample size makes a histogram less noisy; it does not change the underlying model. Compare these views in the normal distribution, then follow counts and waiting times.

Reset all settings