Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Reading a distribution
A random variable assigns a number to an outcome. A distribution tells us how probability is allocated across those numbers. Before interpreting a chart, identify whether the variable is discrete or continuous, and read the vertical axis.
A die has discrete outcomes. Its probability mass function, or PMF, gives the probability of each face. Those six probabilities add to one. A normal model is continuous: its density curve allocates probability through area, not height.
P(-1 ≤ Z ≤ 1) = 0.682689 (68.269%).
CDF at upper bound: 0.841345. Upper tail: 0.158655.
Height is not probability
The chart shows a standard normal density: mean zero and standard deviation one. The height near zero is about 0.399. That does not mean a 39.9% chance of exactly zero. Every individual point has probability zero in this continuous model; an interval can have positive probability.
Set the bounds to −1 and 1. The area between them is about 0.6827. Change them to −2 and 2 and it becomes about 0.9545. Equal lower and upper bounds give zero area.
A density can exceed one. A uniform distribution over an interval of width 0.1 has density 10 throughout that interval, but its total area is still 10 × 0.1 = 1.
Accumulate from the left
The cumulative distribution function, or CDF, is F(x) = P(X ≤ x). It starts near zero and ends near one. For a continuous variable, the probability between a and b is F(b) − F(a).
For an integer count, endpoints require care: P(a ≤ X ≤ b) = F(b) − F(a−1). Subtracting F(a) would accidentally remove the probability at a. The Poisson lesson lets you compare point and cumulative probabilities.
Make a prediction
If F(1.645) is about 0.95 for a standard normal variable, does the interval from −1.645 to 1.645 contain 95%?
Explore the answer
No. Symmetry puts about 5% below −1.645 and 5% above 1.645, leaving about 90% in the middle. A central 95% interval uses approximately ±1.96.
Read in the opposite direction
A quantile asks for a value corresponding to a cumulative probability. The 0.95 quantile of a standard normal distribution is about 1.645. For a discrete distribution, the CDF jumps, so a quantile need not have cumulative probability exactly equal to the requested level: use the smallest value reaching or exceeding it.
Move from probability to a quantile
Requested p = 0.950. Quantile q = 1.644854. P(X<q) = 0.950000; P(X≤q) = 0.950000.
Choose the fair die and request p=0.95. The inverse CDF returns 6, where F(6)=1, because F(5)=5/6 is still below 0.95. The quantile reaches or exceeds the requested probability; it need not equal it.
Now choose the nine observed values and request p=0.25. The inverse empirical CDF returns −0.4: one observation lies strictly below it and three lie at or below it. This convention uses an observed value rather than interpolating between neighboring observations. Software can use other sample-quantile conventions, so state which definition you need.
Make a prediction
Can an empirical CDF jump by more than 1/n at one value?
Explore the answer
Yes. Tied observations share a location, so their probability masses add. Here −0.4 occurs twice and produces a jump of 2/9. Every observation still contributes exactly 1/9.
Separate model and sample
A theoretical curve describes a model. A histogram describes observed or simulated values and changes when you resample. Changing sample size makes a histogram less noisy; it does not change the underlying model. Compare these views in the normal distribution, then follow counts and waiting times.