Updated
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Hypothesis testing
You will learn: Separate observed evidence, effect size, power, and repeated-test error rates.
Start with: Confidence intervals
The frequentist recipe for assessing how incompatible a dataset is with a specified null model:
- State a null hypothesis and an alternative .
- Pick a test statistic whose sampling distribution is known under the null.
- Observe the data, compute the statistic.
- Report the probability of seeing something at least as extreme under the null — the p-value.
Connect the statistic to the measured effect
| Known-σ normal model (σ = 1) | Value |
|---|---|
| Observed sample mean z/√n | 0.357771 |
| 95.0% interval for μ | [-0.0804904, 0.796032] |
| Null mean zero outside interval? | No |
| False rejection probability under μ = 0 | 0.0500000 |
| Rejection probability at chosen true μ | 0.432158 |
This is exact for independent normal observations with known σ = 1 and a sample size chosen before inspecting results. Holding z fixed while changing n changes the implied observed mean. The true μ control describes hypothetical repeated datasets; it does not change this observed p-value. Both density windows extend at least four standard deviations from their centers; calculations retain their full tails.
More chances to reject
If all 20 nulls are true and their tests are independent, testing each at α gives probability 1 − (1 − α)ᵐ = 0.641514 of at least one false rejection. Testing each at α/m = 0.00250000 gives independent-family probability 0.0488301. Bonferroni's at-most-α guarantee also holds without independence by the union bound.
Repeated looks at accumulating data are dependent tests. The independent formula above does not apply to optional stopping. A fixed-sample α cutoff alone does not control the chance of ever crossing it; use a prespecified sequential procedure when stopping depends on results.
What to notice
- P-value = tail area past the observed z. The rose shaded regions are the rejection zone . Drag the observed z-line across the boundary and watch the p-value cross α.
- Power is the alternative’s tail inside the rejection region. The dashed green curve is the distribution under the true μ. Its mass that falls inside the rejection zone is the probability the test detects the effect. Raising moves the alternative mean farther from zero, to , while its variance on the z scale stays one, sending power toward 1 for a fixed nonzero effect.
- This continuous test has uniform p-values under the null. At true μ = 0, the observed z is itself Normal(0, 1), so the tail probability is uniform on [0, 1]. That’s why “p < 0.05 by chance” occurs 5% of the time when the null is true — not a bug, the definition.
What a p-value is not
- Not the probability that H₀ is true. That would require a prior over H₀ (and you’d need Bayes).
- Not the probability that the result is due to chance. “Chance” isn’t a well-defined event.
- Not a measure of effect size. A tiny effect with huge n hits p < 0.001; a huge effect with tiny n can sit at p = 0.3.
Power, effect size, and sample size
For a fixed test, significance level, noise model, and sidedness, effect size and sample size determine power:
Practically:
- Power analysis before the study. Plug in the smallest effect you care about and the power you want (typically 0.8); solve for n.
- Power at the observed effect adds no independent evidence in this model. With fixed α and the same z test, it is determined by the observed |z|. For planning, use a scientifically meaningful effect rather than treating the noisy observed effect as known truth.
State the model and keep the units
The demo assumes independent observations from Normal(μ, 1), known standard deviation one, and a fixed n. Under μ = 0, Z = √n X̄ is exactly standard normal; under a specified μ it is Normal(μ√n, 1). For nonnormal observations the same reference may only be approximate. When the population standard deviation is estimated from normal data, use the Student t procedure instead.
At n = 25 and observed z = 2, the observed mean is 0.4, its standard error is 0.2, and the two-sided p-value is approximately 0.0455003. The 95% interval is approximately [0.008007, 0.791993], excluding zero. The equivalent test/interval comparison requires the same model, confidence level, and sidedness. The demo uses strict p < α and a closed confidence interval; a boundary value is not outside that interval.
Make a prediction
Keep z = 2 and increase n from 25 to 100. Must the p-value decrease?
Explore the answer
No. It remains about 0.0455003 because z is already standardized. The implied observed mean falls from 0.4 to 0.2 and its standard error from 0.2 to 0.1. Holding the observed mean fixed would be a different experiment and would change z.
A family of tests is a different event
For 20 independent true nulls tested at α = 0.05 each, the probability of at least one false rejection is . It is not 5%, and it is not exactly 20 × 5%: the latter sum double-counts outcomes with several rejections. The independence formula is displayed separately from Bonferroni’s union-bound guarantee.
With per-test level α/20, the chance of at least one false rejection is at most α regardless of dependence, provided each individual test is valid and the family is specified. This does not mean that 5% of reported significant results are false, which is a different conditional question.
Make a prediction
Does failing to reject establish that the effect is too small to matter?
Explore the answer
No. The study may be imprecise. Read the effect estimate and interval relative to a meaningful effect range. Demonstrating equivalence or noninferiority requires an explicitly designed hypothesis and decision rule.
Repeated looks are not independent studies
If you inspect accumulating data after every observation and stop at the first p-value below 0.05, the chance of ever crossing the threshold generally differs from the fixed-n false-rejection probability. These looks share data, so the independent-family formula cannot be reused. Prespecified group-sequential boundaries, alpha spending, or other valid sequential methods can address that design. The current demo implements a fixed-sample test and does not claim sequential validity.
Reference and next steps
The ASA statement on p-values discusses model dependence, selective reporting, and why a threshold alone cannot support a scientific conclusion. The z calculations and independent-family probabilities here follow directly from the stated normal model and the complement rule. Continue with confidence intervals, bootstrap uncertainty, and Bayes to compare the questions they answer.