law #15
In this lesson

Wide tables, equations, and code scroll sideways. Swipe, or Tab to focus them and use the left and right arrow keys.

Maximum likelihood

You will learn: Hold observations fixed and distinguish interior, boundary, nonunique, and nonexistent maxima.

Start with: Cross-entropy · Conditioning and independence

Which parameter makes these fixed observations most plausible under a stated model? Maximum likelihood answers that optimization question. Review probability mass versus density, independence, and cross-entropy. Likelihood is the same sampling formula viewed as a function of the parameter after the data have been recorded.

For independent observations with common probability mass or density fθ, the likelihood and its natural logarithm are:

L(θ;x1,…,xn)=∏i=1nfθ(xi),ℓ(θ)=∑i=1nlog⁡fθ(xi).L(\theta;x_1,\ldots,x_n)=\prod_{i=1}^n f_\theta(x_i),\qquad \ell(\theta)=\sum_{i=1}^n\log f_\theta(x_i).

The logarithm preserves the maximizer and avoids multiplying many tiny values. Independence justifies the product here; dependent observations require their joint model. For continuous observations the factors are densities, not probabilities of the exact measured values.

These counts specify a synthetic fixed sample. Changing candidate p keeps the observations fixed. Both zero means no observations, so every p maximizes the constant likelihood.

Likelihood divided by its maximum, with the observed sample held fixed00.20.40.60.8100.20.40.60.81candidate prelative likelihood

Blue: relative likelihood; open red point: candidate; dashed line: maximizer when unique. The peak is one by division, not because the curve is a probability distribution over parameters.

Fixed-sample quantityValue
Observations n10
Maximum-likelihood estimate0.7000000
Candidate log-likelihood-6.931472
Log-likelihood minus maximum-0.8228288
Candidate relative likelihood0.4391875

Natural logarithms are used. Impossible observed data give log-likelihood −∞ and relative likelihood zero. Very small positive relative likelihoods may underflow to zero in floating-point arithmetic; the finite log-likelihood still distinguishes them. No confidence or posterior probability is assigned by this ratio.

Compare parameter values for one stated observation model and a fixed sample. A different data-recording protocol, dependence structure, or prior defines a different inference problem.

A complete Bernoulli calculation

With h successes and t failures from independent Bernoulli trials, the likelihood of a particular observed sequence is pʰ(1−p)ᵗ. Grouping the observations into a count multiplies this by a binomial coefficient that is constant in p, so the maximizing p is unchanged under this fixed-trial model.

The default h=7,t=3 gives n=10. At p=1/2 the likelihood is (1/2)¹⁰≈0.000976563. At p=0.7 it is 0.7⁷·0.3³≈0.002223566. Their ratio is about 0.439188, meaning the candidate likelihood is 43.9% of the maximum. It is not a 43.9% probability that p=1/2 is correct.

For 0<p<1, differentiation gives:

ℓ′(p)=hp−t1−p,ℓ′′(p)=−hp2−t(1−p)2.\ell'(p)=\frac{h}{p}-\frac{t}{1-p},\qquad \ell''(p)=-\frac{h}{p^2}-\frac{t}{(1-p)^2}.

When both counts are positive, the second derivative is negative and the stationary point p̂=h/(h+t) is the unique maximum. Moving candidate p changes the likelihood curve’s inspected point while leaving the data unchanged. Changing a count poses a new inference problem.

If every observation is a success, the maximum on the closed parameter space [0,1] is p̂=1; if every observation is a failure, it is p̂=0. An interior zero-gradient search would miss these cases. With no observations, every p has likelihood one and there is no unique estimate. On an open parameter space (0,1), an all-success sample has a supremum as p approaches one, but no attained maximum.

Make a prediction

Do seven successes guarantee the true success probability is 0.7?

Explore the answer

Only seven successes together with three failures produce this particular estimate. Even then, the estimate is a function of one sample. Repeated samples can produce other estimates, and the model assumptions may be wrong. Maximizing likelihood does not remove sampling uncertainty.

A maximum where the derivative is not zero

Switch to independent Uniform(0,θ) observations, with θ>0 and density 1/θ for 0≤x≤θ. Let m be the largest observation. For a nonnegative sample:

L(θ)=θ−n1{θ≥m},θ^=m(m>0).L(\theta)=\theta^{-n}\mathbf 1\{\theta\ge m\},\qquad \hat\theta=m\quad(m>0).

For the sample 1,2,3, any θ<3 assigns zero density to at least one observation. At θ=3 the likelihood is 1/27; at θ=4 it is 1/64. The ratio is 27/64=0.421875. Above three, the likelihood decreases, so the maximum sits at the support boundary even though the derivative there does not vanish.

If all observed values are zero, θ⁻ⁿ grows without bound as θ approaches zero from above. No positive θ maximizes it. The experiment explicitly removes the normalized curve in this case. The stated inclusive endpoint convention matters for an attained maximum; changing the density’s support convention can change attainment even when it describes the same continuous probability law.

The uniform MLE is also biased: under the model, P(m≤z)=(z/θ)ⁿ for 0≤z≤θ. Integrating its tail gives E[m]=nθ/(n+1), below θ. Multiplying m by (n+1)/n removes that bias in this model, but then the result is no longer the MLE. Optimization and unbiasedness are different properties.

Likelihood, posterior, and uncertainty

The graph divides likelihood by its maximum, so its peak is one. It does not normalize its area. A Bayesian update instead multiplies by a prior and normalizes the resulting posterior when possible. A continuous parameter can have a posterior density, while its exact point probabilities remain zero.

For a fixed empirical categorical distribution, maximizing independent-sample likelihood is equivalent to minimizing cross-entropy. This does not make a training score a test of generalization, or guarantee that a complicated likelihood has one well-behaved maximum. Familiar large-sample approximations need conditions such as identification and suitable regularity; parameter-dependent support and boundary estimates deserve separate analysis.

Make a prediction

Does a relative likelihood above 0.05 define a 95% confidence interval?

Explore the answer

No. A likelihood ratio is not automatically a coverage probability. A confidence procedure needs its own sampling-distribution justification. This experiment reports the ratio without assigning it a confidence level.

Continue to confidence intervals for repeated-sample coverage, or conjugate priors for posterior and predictive calculations.

Sources

StatLect’s maximum-likelihood introduction explains the fixed-data optimization and conditions behind asymptotic results. The Bernoulli derivatives, support-boundary calculation, and uniform-maximum expectation above are derived explicitly; they can be checked without a numerical optimizer.

Reset all settings