tokens assigned the heuristic accusative label among
n=12{,}280
pronoun observations. Under an IID Bernoulli model, which value of the heuristic-label probability \pi makes these data most probable?
The populations and samples page introduced an estimate as a numerical value computed from an observed sample. Here the data are the observed case labels, the model is a Bernoulli distribution with a shared probability \pi, and maximum likelihood supplies the rule used to obtain the estimate. We answer the opening question in two passes: first hold \pi fixed and vary the possible data; then hold the observed data fixed and compare values of \pi.
Holding the parameter fixed first
For one observation,
p(x_i\mid\pi)
=\pi^{x_i}(1-\pi)^{1-x_i}.
If we fix \pi=.30, this PMF assigns probability .30 to an accusative token and .70 to a nonaccusative token. It is a probability distribution over possible values of x_i.
Suppose the first four observations are 1,0,0,1. With \pi=.30, the model assigns this particular sequence probability
(.30)(.70)(.70)(.30)=.0441.
If we instead fix \pi=.50, the probability is
(.50)(.50)(.50)(.50)=.0625.
For this tiny sequence, .50 assigns more probability than .30. The complete sample contains much more information, but the comparison follows the same logic.
Collecting the k observations with x_i=1 and the n-k observations with x_i=0 yields
p(\mathbf{x}\mid\pi)
=\pi^k(1-\pi)^{n-k}.
Holding the data fixed
The likelihood function uses the same expression but treats the observed data as fixed and varies the parameter:
L(\pi;\mathbf{x})
=\pi^{3327}(1-\pi)^{8953}.
We are no longer asking which data values might occur for a fixed \pi. We are comparing possible values of \pi for the data already observed.
The likelihood does not need to sum or integrate to one over \pi. It is not a probability distribution on the parameter. The later page on posterior distributions explains what must be added before probability is assigned to an unknown parameter.
That distinction can be obscured because the likelihood and PMF use the same formula. The distinction lies in what is allowed to vary.
expression
held fixed
allowed to vary
purpose
p(x\mid\pi)
\pi
possible data x
describe a distribution over observations
L(\pi;\mathbf{x})
observed \mathbf{x}
candidate parameter \pi
compare parameter values for those observations
Likelihood values have a relative interpretation. If L(.27;\mathbf{x}) is larger than L(.20;\mathbf{x}), the observed sample is more probable under the first candidate than under the second. This comparison does not produce \mathbb{P}(\pi=.27\mid\mathbf{x}). A probability distribution over \pi requires the prior distribution and normalization introduced later.
The likelihood can also be written for the count K=k rather than the complete ordered sequence. The binomial likelihood then contains the factor {n\choose k}. That factor depends on the observed data but not on \pi, so it changes the likelihood’s height without changing its maximizing value. The data representation must still be stated because the two likelihood values are not numerically identical.
Using the log-likelihood
Multiplying 12,280 probabilities produces extremely small numbers. Taking logarithms turns the product into a sum:
\ell(\pi;\mathbf{x})
=k\log\pi+(n-k)\log(1-\pi).
The logarithm is increasing, so the likelihood and log-likelihood reach their maxima at the same parameter value.
We can compare two candidates without evaluating very small products. For the full corpus extract, the difference
\ell(.27;\mathbf{x})-\ell(.20;\mathbf{x})
is positive. Thus .27 assigns greater likelihood to the observations than .20. Taking logs changes the numerical scale, but it does not reverse the ordering of candidate parameter values.
We can inspect the curve on a grid:
Code
k <-3327n <-12280pi_grid <-seq(.20, .35, length.out =300)log_likelihood <- k *log(pi_grid) + (n - k) *log(1- pi_grid)plot(pi_grid, log_likelihood, type ="l",xlab =expression(pi), ylab ="log-likelihood")abline(v = k / n, lty =2)
Finding the maximum
Now we need the value at the top of the curve. Differentiating and setting the derivative to zero gives
The notation arg max returns an argument, not a height. Here the argument is a candidate value of \pi. The maximum likelihood is the value of the function at that argument. In ordinary discussion, these two objects are sometimes conflated, but the estimate is the parameter value.
A point where the first derivative equals zero is called a stationary point. Setting this expression to zero gives \widehat{\pi}=k/n. To establish that this stationary point is a maximum, inspect the second derivative:
For 0<\pi<1, both terms are negative. Thus the log-likelihood is concave over its interior, and its stationary point is the unique interior maximum.
Interpreting only what was estimated
The estimate .271 is the observed proportion of tokens assigned the heuristic accusative label in this extract. Under the IID Bernoulli model, it estimates one shared label probability for the population represented by the sampling process. It does not say that every pronoun type, speaker, document, or syntactic context has an accusative probability of .271.
This limitation is a model limitation. Pronoun case depends on grammatical function, lexical form, and construction type. If those sources of variation matter to the scientific question, a one-parameter Bernoulli model collapses them. Maximum likelihood can select the best value of \pi within that simple family, but it cannot make the family express distinctions it does not contain.
The calculation has produced one point estimate under one sampling model. The next page applies the same rule to a count model. Properties of estimators then asks how an estimation rule behaves across many possible samples.
NoteCheck the roles
Explain the difference between p(x\mid\pi) and L(\pi;x) by stating which object is held fixed and which is allowed to vary in each expression.
For observations 1,0,1,1,0, calculate the maximum-likelihood estimate of \pi.
Explain why taking a logarithm does not move the maximizing value of \pi.
Identify one linguistic source of variation that the shared Bernoulli probability omits in the pronoun case example.
---title: "Maximum likelihood"---The corpus extract contains$$k=3{,}327$$tokens assigned the heuristic accusative label among$$n=12{,}280$$pronoun observations. Under an IID Bernoulli model, which value of the heuristic-label probability $\pi$ makes these data most probable?The [populations and samples page](populations-and-samples.qmd) introduced an estimate as a numerical value computed from an observed sample. Here the [**data**](https://openstax.org/books/introductory-statistics-2e/pages/1-1-definitions-of-statistics-probability-and-key-terms) are the observed case labels, the [**model**](https://openstax.org/books/introductory-statistics-2e/pages/1-1-definitions-of-statistics-probability-and-key-terms) is a Bernoulli distribution with a shared probability $\pi$, and maximum likelihood supplies the rule used to obtain the estimate. We answer the opening question in two passes: first hold $\pi$ fixed and vary the possible data; then hold the observed data fixed and compare values of $\pi$.## Holding the parameter fixed firstFor one observation,$$p(x_i\mid\pi)=\pi^{x_i}(1-\pi)^{1-x_i}.$$If we fix $\pi=.30$, this PMF assigns probability $.30$ to an accusative token and $.70$ to a nonaccusative token. It is a probability distribution over possible values of $x_i$.Suppose the first four observations are $1,0,0,1$. With $\pi=.30$, the modelassigns this particular sequence probability$$(.30)(.70)(.70)(.30)=.0441.$$If we instead fix $\pi=.50$, the probability is$$(.50)(.50)(.50)(.50)=.0625.$$For this tiny sequence, $.50$ assigns more probability than $.30$. The completesample contains much more information, but the comparison follows the samelogic.For the complete sample, independence gives$$p(\mathbf{x}\mid\pi)=\prod_{i=1}^{n}\pi^{x_i}(1-\pi)^{1-x_i}.$$Collecting the $k$ observations with $x_i=1$ and the $n-k$ observations with $x_i=0$ yields$$p(\mathbf{x}\mid\pi)=\pi^k(1-\pi)^{n-k}.$$## Holding the data fixedThe [**likelihood function**](https://bruno.nicenboim.me/bayescogsci/ch-intro.html) uses the same expression but treats the observed data as fixed and varies the parameter:$$L(\pi;\mathbf{x})=\pi^{3327}(1-\pi)^{8953}.$$We are no longer asking which data values might occur for a fixed $\pi$. We are comparing possible values of $\pi$ for the data already observed.The likelihood does not need to sum or integrate to one over $\pi$. It is not a probability distribution on the parameter. The later page on [posterior distributions](posterior-distributions.qmd) explains what must be added before probability is assigned to an unknown parameter.That distinction can be obscured because the likelihood and PMF use the sameformula. The distinction lies in what is allowed to vary.| expression | held fixed | allowed to vary | purpose ||---|---|---|---|| $p(x\mid\pi)$ | $\pi$ | possible data $x$ | describe a distribution over observations || $L(\pi;\mathbf{x})$ | observed $\mathbf{x}$ | candidate parameter $\pi$ | compare parameter values for those observations |Likelihood values have a relative interpretation. If$L(.27;\mathbf{x})$ is larger than $L(.20;\mathbf{x})$, the observed sample ismore probable under the first candidate than under the second. This comparisondoes not produce $\mathbb{P}(\pi=.27\mid\mathbf{x})$. A probability distributionover $\pi$ requires the prior distribution and normalization introduced later.The likelihood can also be written for the count $K=k$ rather than the complete ordered sequence. The binomial likelihood then contains the factor ${n\choose k}$. That factor depends on the observed data but not on $\pi$, so it changes the likelihood's height without changing its maximizing value. The data representation must still be stated because the two likelihood values are not numerically identical.## Using the log-likelihoodMultiplying 12,280 probabilities produces extremely small numbers. Taking logarithms turns the product into a sum:$$\ell(\pi;\mathbf{x})=k\log\pi+(n-k)\log(1-\pi).$$The logarithm is increasing, so the likelihood and log-likelihood reach their maxima at the same parameter value.We can compare two candidates without evaluating very small products. For thefull corpus extract, the difference$$\ell(.27;\mathbf{x})-\ell(.20;\mathbf{x})$$is positive. Thus $.27$ assigns greater likelihood to the observations than$.20$. Taking logs changes the numerical scale, but it does not reverse theordering of candidate parameter values.We can inspect the curve on a grid:```{r}#| label: pronoun-likelihood#| echo: truek <-3327n <-12280pi_grid <-seq(.20, .35, length.out =300)log_likelihood <- k *log(pi_grid) + (n - k) *log(1- pi_grid)plot(pi_grid, log_likelihood, type ="l",xlab =expression(pi), ylab ="log-likelihood")abline(v = k / n, lty =2)```## Finding the maximumNow we need the value at the top of the curve. Differentiating and setting the derivative to zero gives$$\widehat{\pi}_{\mathrm{MLE}}=\frac{k}{n}.$$For this extract,$$\widehat{\pi}_{\mathrm{MLE}}=\frac{3327}{12280}\approx.271.$$The [**maximum-likelihood estimate**](https://bruno.nicenboim.me/bayescogsci/ch-intro.html) is the parameter value that maximizes the likelihood:$$\widehat{\theta}_{\mathrm{MLE}}=\operatorname*{arg\,max}_{\theta}L(\theta;\mathbf{x}).$$The notation `arg max` returns an argument, not a height. Here the argument is acandidate value of $\pi$. The maximum likelihood is the value of the function atthat argument. In ordinary discussion, these two objects are sometimesconflated, but the estimate is the parameter value.## Checking the stationary pointThe derivative of the Bernoulli log-likelihood is$$\frac{\mathrm{d}}{\mathrm{d}\pi}\ell(\pi;\mathbf{x})=\frac{k}{\pi}-\frac{n-k}{1-\pi}.$$A point where the first derivative equals zero is called a [**stationary point**](https://mathworld.wolfram.com/StationaryPoint.html).Setting this expression to zero gives $\widehat{\pi}=k/n$. To establish that thisstationary point is a maximum, inspect the second derivative:$$\frac{\mathrm{d}^2}{\mathrm{d}\pi^2}\ell(\pi;\mathbf{x})=-\frac{k}{\pi^2}-\frac{n-k}{(1-\pi)^2}.$$For $0<\pi<1$, both terms are negative. Thus the log-likelihood is concave overits interior, and its stationary point is the unique interior maximum.## Interpreting only what was estimatedThe estimate $.271$ is the observed proportion of tokens assigned the heuristic accusative label in thisextract. Under the IID Bernoulli model, it estimates one shared label probability forthe population represented by the sampling process. It does not say that everypronoun type, speaker, document, or syntactic context has an accusativeprobability of $.271$.This limitation is a model limitation. Pronoun case depends on grammaticalfunction, lexical form, and construction type. If those sources of variationmatter to the scientific question, a one-parameter Bernoulli model collapsesthem. Maximum likelihood can select the best value of $\pi$ within that simplefamily, but it cannot make the family express distinctions it does not contain.The calculation has produced one point estimate under one sampling model. The [next page](poisson-maximum-likelihood.qmd) applies the same rule to a count model. [Properties of estimators](properties-of-estimators.qmd) then asks how an estimation rule behaves across many possible samples.::: {.callout-note title="Check the roles"}1. Explain the difference between $p(x\mid\pi)$ and $L(\pi;x)$ by stating which object is held fixed and which is allowed to vary in each expression.2. For observations $1,0,1,1,0$, calculate the maximum-likelihood estimate of $\pi$.3. Explain why taking a logarithm does not move the maximizing value of $\pi$.4. Identify one linguistic source of variation that the shared Bernoulli probability omits in the pronoun case example.:::