Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Logistic distribution
You will learn: Convert between cumulative probabilities, quantiles, and odds.
Start with: Reading a distribution
The Normal’s slightly heavier-tailed cousin. Its CDF is the sigmoid that binary logistic regression uses to map a linear predictor to a probability:
The corresponding density:
Move between values, probabilities, and odds
The original cutoff is x=1.00000. P(X≤x) = 0.731059 and P(X>x) = 0.268941. Cumulative odds are exp((x−μ)/s) = 2.71828; increasing x by s multiplies these odds by e.
The 75% quantile is 1.098612. Scale s=1 gives SD 1.81380 and variance 3.28987. The middle 50% spans μ±s log(3).
1/1000 samples fall outside the density window. Exact omitted probability is 0.000670700; both CDF tails use the full distribution.
The regression connection has a different outcome
A logistic random variable is continuous. Binary logistic regression models a Bernoulli outcome with p=sigmoid(η), where η is a predictor. It does not assume the observed zero/one outcome has this continuous density. The CDF above and that inverse-link function share the same formula.
What to notice
- Shape close to Normal, tails a touch heavier. The matched-variance Normal is shorter at the centre and disappears faster in the tails. Kurtosis of 4.2 vs. 3 for the Normal.
- Scale parameter s is not the standard deviation. The variance is , so an s of 1 corresponds to a standard deviation of about 1.81.
- Closed-form quantile. Unlike the Normal, you can invert the CDF with elementary functions: — which is exactly how this demo samples from it.
Why it matters
The difference of two independent standard maximum-type Gumbel variables is standard logistic. Independence and the shared scale matter. This supplies a random-utility construction of the binary logit probability; sharing a formula does not imply that every use of a sigmoid has that generative interpretation.
Calculate a cutoff and invert it
With μ=2 and s=3, x=5 has standardized coordinate (5−2)/3=1. Its cumulative probability is 1/(1+e⁻¹)≈0.731059. The upper tail is about 0.268941. These add to one; a point probability at x=5 is zero.
For a 75th percentile, solve p/(1−p)=exp((x−μ)/s). This gives x=μ+s log(3)=5.295837. The lower quartile is μ−s log(3), so the interquartile range is 2s log(3). Increasing μ translates both cutoffs, while increasing s widens their separation.
Scale s=1 does not mean unit variance. The SD is π/√3≈1.813799. To obtain unit SD, use s=√3/π≈0.551329. A matched-variance normal has a lower central density than logistic, even though logistic also has heavier far tails. The normalization used for a shape comparison must be stated.
Turn a continuous threshold into a binary response
Suppose ε is standard logistic and define a binary response Y=1 when η+ε>0. Symmetry gives P(Y=1)=P(ε>−η)=sigmoid(η). This is one latent-error construction of logistic regression. Conditional on a predictor, Y itself remains Bernoulli; the continuous error ε and observed response Y have different distributions.
Changing η by one multiplies the odds by e. It does not add a fixed amount to the probability: moving η from 0 to 1 changes p from 0.5 to about 0.7311, while moving from 3 to 4 changes p from about 0.9526 to 0.9820.
Make a prediction
If a fitted binary probability is 0.75, does that mean the observed response is a logistic random variable with mean 0.75?
Explore the answer
No. The response is zero or one with probabilities 0.25 and 0.75 in the Bernoulli model. The corresponding log odds are log(3). The logistic CDF supplies the link between the predictor and that probability.
References
Random Services: logistic distribution derives its CDF, inverse, moments, and related constructions. Compare Bernoulli indicators for the binary response and Gumbel maxima for the connected continuous family.