Published
In this lesson
Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.
Maximum likelihood
You will learn: Hold observations fixed and distinguish interior, boundary, nonunique, and nonexistent maxima.
Start with: Cross-entropy · Conditioning and independence
Which parameter makes these fixed observations most plausible under a stated model? Maximum likelihood answers that optimization question. Review probability mass versus density, independence, and cross-entropy. Likelihood is the same sampling formula viewed as a function of the parameter after the data have been recorded.
For independent observations with common probability mass or density fθ, the likelihood and its natural logarithm are:
The logarithm preserves the maximizer and avoids multiplying many tiny values. Independence justifies the product here; dependent observations require their joint model. For continuous observations the factors are densities, not probabilities of the exact measured values.
These counts specify a synthetic fixed sample. Changing candidate p keeps the observations fixed. Both zero means no observations, so every p maximizes the constant likelihood.
Blue: relative likelihood; open red point: candidate; dashed line: maximizer when unique. The peak is one by division, not because the curve is a probability distribution over parameters.
| Fixed-sample quantity | Value |
|---|---|
| Observations n | 10 |
| Maximum-likelihood estimate | 0.7000000 |
| Candidate log-likelihood | -6.931472 |
| Log-likelihood minus maximum | -0.8228288 |
| Candidate relative likelihood | 0.4391875 |
Natural logarithms are used. Impossible observed data give log-likelihood −∞ and relative likelihood zero. Very small positive relative likelihoods may underflow to zero in floating-point arithmetic; the finite log-likelihood still distinguishes them. No confidence or posterior probability is assigned by this ratio.
A complete Bernoulli calculation
With h successes and t failures from independent Bernoulli trials, the likelihood of a particular observed sequence is pʰ(1−p)ᵗ. Grouping the observations into a count multiplies this by a binomial coefficient that is constant in p, so the maximizing p is unchanged under this fixed-trial model.
The default h=7,t=3 gives n=10. At p=1/2 the likelihood is (1/2)¹⁰≈0.000976563. At p=0.7 it is 0.7⁷·0.3³≈0.002223566. Their ratio is about 0.439188, meaning the candidate likelihood is 43.9% of the maximum. It is not a 43.9% probability that p=1/2 is correct.
For 0<p<1, differentiation gives:
When both counts are positive, the second derivative is negative and the stationary point p̂=h/(h+t) is the unique maximum. Moving candidate p changes the likelihood curve’s inspected point while leaving the data unchanged. Changing a count poses a new inference problem.
If every observation is a success, the maximum on the closed parameter space [0,1] is p̂=1; if every observation is a failure, it is p̂=0. An interior zero-gradient search would miss these cases. With no observations, every p has likelihood one and there is no unique estimate. On an open parameter space (0,1), an all-success sample has a supremum as p approaches one, but no attained maximum.
Make a prediction
Do seven successes guarantee the true success probability is 0.7?
Explore the answer
Only seven successes together with three failures produce this particular estimate. Even then, the estimate is a function of one sample. Repeated samples can produce other estimates, and the model assumptions may be wrong. Maximizing likelihood does not remove sampling uncertainty.
A maximum where the derivative is not zero
Switch to independent Uniform(0,θ) observations, with θ>0 and density 1/θ for 0≤x≤θ. Let m be the largest observation. For a nonnegative sample:
For the sample 1,2,3, any θ<3 assigns zero density to at least one observation. At θ=3 the likelihood is 1/27; at θ=4 it is 1/64. The ratio is 27/64=0.421875. Above three, the likelihood decreases, so the maximum sits at the support boundary even though the derivative there does not vanish.
If all observed values are zero, θ⁻ⁿ grows without bound as θ approaches zero from above. No positive θ maximizes it. The experiment explicitly removes the normalized curve in this case. The stated inclusive endpoint convention matters for an attained maximum; changing the density’s support convention can change attainment even when it describes the same continuous probability law.
The uniform MLE is also biased: under the model, P(m≤z)=(z/θ)ⁿ for 0≤z≤θ. Integrating its tail gives E[m]=nθ/(n+1), below θ. Multiplying m by (n+1)/n removes that bias in this model, but then the result is no longer the MLE. Optimization and unbiasedness are different properties.
Likelihood, posterior, and uncertainty
The graph divides likelihood by its maximum, so its peak is one. It does not normalize its area. A Bayesian update instead multiplies by a prior and normalizes the resulting posterior when possible. A continuous parameter can have a posterior density, while its exact point probabilities remain zero.
For a fixed empirical categorical distribution, maximizing independent-sample likelihood is equivalent to minimizing cross-entropy. This does not make a training score a test of generalization, or guarantee that a complicated likelihood has one well-behaved maximum. Familiar large-sample approximations need conditions such as identification and suitable regularity; parameter-dependent support and boundary estimates deserve separate analysis.
Make a prediction
Does a relative likelihood above 0.05 define a 95% confidence interval?
Explore the answer
No. A likelihood ratio is not automatically a coverage probability. A confidence procedure needs its own sampling-distribution justification. This experiment reports the ratio without assigning it a confidence level.
Continue to confidence intervals for repeated-sample coverage, or conjugate priors for posterior and predictive calculations.
Sources
StatLect’s maximum-likelihood introduction explains the fixed-data optimization and conditions behind asymptotic results. The Bernoulli derivatives, support-boundary calculation, and uniform-maximum expectation above are derived explicitly; they can be checked without a numerical optimizer.