Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Bernoulli distribution
You will learn: Represent a two-outcome trial and connect its probability, mean, and variance.
Start with: Reading a distribution
A single trial with two outcomes: 1 with probability , 0 otherwise. Binomial counts and geometric or negative-binomial waiting times can be built from independent Bernoulli trials with a common success probability.
One event, many repetitions
Trial 1: the generator draws U=0.627074. The indicator 1(U<p) is 0. Changing p keeps these underlying uniform draws fixed.
| Quantity | Value |
|---|---|
| Model probability and mean p | 0.3000 |
| One-trial variance p(1−p) | 0.2100 |
| Observed successes / repetitions | 72 / 200 = 0.3600 |
| SD of a proportion from n independent trials | 0.03240 |
An individual indicator is always zero or one, even when p is fractional. The observed proportion estimates p across repetitions; its error can increase when another draw is added. The displayed standard deviation uses the model's known p, not an estimate or a confidence interval.
What to notice
- Only one knob. The whole distribution is determined by . The dark bars show the exact PMF; the light bars show the empirical frequencies. As n grows, empirical proportions converge almost surely under independent common-p repetitions — that’s the Law of Large Numbers.
- Variance is maximized at p = 0.5, where you’re least sure of the outcome, and shrinks to zero at p = 0 or p = 1.
- Summing n independent Bernoullis gives a Binomial(n, p). The common success probability is essential; unequal probabilities give a different count distribution.
Why it matters
The Bernoulli is the indicator function for any yes/no event. Expressing as “the probability equals the expected value of the indicator” is the move that underlies most elementary probability identities.
An indicator can represent any event
Roll a fair six-sided die and let X=1 when the result exceeds four, otherwise X=0. The die has six outcomes, but this indicator has two: p=2/6=1/3. Therefore E[X]=1/3 and Var(X)=2/9. An expected value of one third does not mean the observed indicator ever takes that value.
Since X²=X, its variance is E[X²]−E[X]²=p−p². Add ten such indicators from independent die rolls: the expected total is 10/3 and its variance is 20/9. Dividing the total by ten gives a proportion with variance 1/45. The single-trial distribution stays Bernoulli as repetitions increase; only the estimate becomes more precise.
Known p and an estimated proportion answer different questions
The slider specifies p for the synthetic generator. The displayed success fraction is computed from the generated sample. At p=0.3 and n=200, the model standard deviation of this fraction is √(0.21/200)≈0.032404. A different seed can move the estimate in either direction. Increasing n is not a promise that every new estimate gets closer.
At p=0 or p=1, the outcome is certain and both variances are zero. In observed data, seeing no successes in a small sample does not establish that the underlying p is zero. Estimating an unknown p requires an uncertainty model, as the beta lesson illustrates.
Make a prediction
You copy one die roll's indicator ten times. Is the variance of their average 1/45?
Explore the answer
No. The average equals that one indicator, with variance 2/9. Its marginal distribution is still Bernoulli, but copying does not create independent repetitions.
Reference and next step
Random Services: Bernoulli trials derives indicator moments and the uniform-threshold construction. Continue to binomial counts to distinguish a single event, its repeated count, and the normal approximation.