Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Negative binomial distribution
You will learn: Distinguish trials from failures while waiting for a specified number of successes.
Start with: Geometric distribution
Flip a biased coin with success probability until you accumulate exactly successes. The number of failures along the way is negative-binomial distributed.
Total trials equal failures + 3. The distribution shifts by 3; its variance stays 11.250. Mean under this convention: 4.500.
P(failures ≤ 5) = 0.684605. P(exactly 5) = 0.104509.
1 of 1000 simulated counts exceed the chart window; they remain in the frequency denominator. The exact probability calculation includes the entire support.
What to notice
- Raising shifts mass to the right (you’re waiting for more successes, so more failures pile up) and smooths the shape toward a Normal, since the total is a sum of independent geometric waits.
- Lowering stretches the whole distribution — rare successes mean long waits.
- Variance exceeds the mean by a factor of . Compared to Poisson (where variance equals the mean), the negative binomial is overdispersed — which is why it’s the go-to for count data that clumps.
Relationship to other distributions
- Geometric. With , you’re waiting for the first success, which is the failure-count geometric distribution. Our geometric page instead counts the successful trial too, so its values are one larger.
- Sum of geometrics. More generally, a negative binomial with parameter is the sum of independent geometrics — using the failure-count convention. This parallels a positive-integer-shape gamma distribution, built from independent common-rate exponential waits.
- Poisson-gamma mixture. A Poisson whose rate is itself gamma-distributed integrates out to a negative binomial. This relationship requires gamma variation in the Poisson mean; arbitrary variation need not give a negative binomial.
Two counters for one experiment
At r=3 and p=0.4, suppose the sequence is failure, success, failure, failure, success, success. There are three failures and six total trials. Switching the counter adds r to every result; it changes the mean from r(1−p)/p=4.5 to r/p=7.5 but leaves the variance at r(1−p)/p²=11.25.
Exactly two failures before the third success requires the fifth trial to succeed and exactly two of the first four to succeed. There are six such arrangements, each with probability 0.4³×0.6², giving 0.13824. The compute panel adds these masses to answer an at-most question under either convention. A total-trial threshold below r is impossible.
Why varying Poisson means produce overdispersion
Suppose a count conditional on Λ is Poisson(Λ), and Λ has a gamma distribution with shape r and rate β. Integrating over Λ gives a failure-count negative binomial with p=β/(β+1). Conversely β=p/(1−p). Its mean is r/β and variance is r/β+r/β²: variation in Λ adds the second term. This is an exact model relationship, not a proof that any overdispersed dataset was generated this way.
At r=3 and p=0.4, β=2/3. The average conditional Poisson mean is 4.5, but the marginal variance is 11.25. A fixed-rate Poisson with mean 4.5 has variance 4.5 instead. The stopping experiment uses integer r; the gamma-mixture family also permits positive noninteger shapes.
Make a prediction
A library's geometric function counts trials starting at one. Can you sum r such draws and call that the failure count?
Explore the answer
Subtract r first. The sum includes one successful trial for each wait. Its unshifted value is the total-trial negative binomial, while this page’s default counts only failures.
References
Random Services: negative binomial uses the total-trial convention; Stanford: mixture models derives the gamma–Poisson mixture and its parameter conversion.