The maximum-likelihood page introduced estimation with a Bernoulli response. We now repeat the same reasoning with the Poisson distribution for a count response. The question is whether maximizing a count likelihood leads to a familiar sample summary. Deriving the answer will also show which steps depend specifically on the Poisson model.
Starting with a linguistic count
The Universal Dependencies English Web Treebank contains annotated dependency relations. Consider a legacy teaching extract of 2,594 candidate wh-form tokens and the linear distance from each selected token to its annotated head. A relative pronoun such as who may be near its head, as in the officer who was working, or separated from it by several intervening tokens.
NoteData source
The extract comes from the Universal Dependencies English Web Treebank, derived from the English Web Treebank described by Silveira et al. (2014). The variable used here is the absolute difference between a selected token’s index and its head’s index in the basic dependency annotation. The stored file retains legacy column names such as gap_position, but that position is the annotated head token, not a silent gap. The selection heuristic also admitted some tokens without a UD PronType value, including instances of however. Thus the rows are candidate wh-form tokens, not a gold-standard set of wh dependencies. The distance is linear and annotation-dependent; it is not a filler-gap distance or a measure of syntactic depth.
Let X_i be this selected-token-to-head distance for observation i. The observed distances are positive integers, but the Poisson family also assigns probability to zero. We will temporarily ignore that mismatch so that we can isolate the estimation step. Our working model is
The model says that every observation shares one Poisson mean \lambda and that the observations are independent after we condition on that mean. Those claims are strong. We will check one of their consequences after finding the estimate.
Writing the likelihood one observation at a time
For one observed distance x_i, the Poisson probability mass is
This is the likelihood function. It compares candidate values of the mean by asking how much probability each candidate assigns to the observed collection of distances.
Taking logs before differentiating
The product contains one small probability for each observation. Taking a log turns that product into a sum:
The final term depends on the observations but not on \lambda. It changes the height of the log-likelihood curve, but it cannot change where that curve is maximized.
The maximum-likelihood estimate of a shared Poisson mean is the sample mean. This equality is a consequence of the Poisson model and its likelihood. It is not a general definition of maximum likelihood.
Verifying the result on the dependency distances
The data are available in wh_filler_gap.csv. We can compute the analytic estimate and compare it with a numerical optimizer.
Both calculations return approximately 3.007 tokens. The agreement is not a second piece of evidence about the linguistic phenomenon. It is a check that the code and the algebra solve the same optimization problem.
We can also inspect the log-likelihood near the maximum.
The curve is highest at the sample mean. Values near the maximum have similar support from the observed data, while more distant values assign lower joint probability to those observations.
Separating estimation from model adequacy
The estimate answers a conditional question: if the distances were independent draws from one Poisson distribution, which value of \lambda would fit them best? It does not establish that the Poisson family is a good description.
A Poisson variable has equal mean and variance. The fitted model must thus have both mean and variance near 3.01, while the empirical variance is more than three times as large. The distances also cannot be zero in this particular extract, even though the fitted distribution assigns positive probability to zero.
These discrepancies identify two distinct model failures: the support mismatch concerns which values can occur, while the variance mismatch concerns how probability is distributed over the values that can occur. Maximum likelihood finds the best parameter within the chosen family. It cannot repair a family whose structural assumptions miss the response.
NoteWork through the estimate
For the distances 1,2,2,3,7, calculate \widehat{\lambda}_{\mathrm{MLE}}.
Explain why \sum_i\log(x_i!) disappears when we differentiate with respect to \lambda.
State the conditional question answered by the estimate.
Give one reason that selected-token-to-head distances from the same document might not be independent.
---title: "Maximum likelihood for a Poisson mean"---The [maximum-likelihood page](maximum-likelihood.qmd) introduced estimation witha Bernoulli response. We now repeat the same reasoning with the [Poisson distribution](../random-variables-and-distributions/poisson-distribution.qmd) for a count response. The question is whether maximizing a count likelihood leads to a familiar sample summary. Deriving the answer will also show which steps depend specifically on the Poisson model.## Starting with a linguistic countThe Universal Dependencies English Web Treebank contains annotated dependencyrelations. Consider a legacy teaching extract of 2,594 candidate wh-form tokensand the linear distance from each selected token to its annotated head. Arelative pronoun such as *who* may be near its head, as in *the officer who wasworking*, or separated from it by several intervening tokens.::: {.callout-note title="Data source"}The extract comes from the [Universal Dependencies English WebTreebank](https://universaldependencies.org/treebanks/en_ewt/index.html), derivedfrom the English Web Treebank described by [Silveira et al.(2014)](https://aclanthology.org/L14-1067/). The variable used here is theabsolute difference between a selected token's index and its head's index in thebasic dependency annotation. The stored file retains legacy column names suchas `gap_position`, but that position is the annotated head token, not a silentgap. The selection heuristic also admitted some tokens without a UD `PronType`value, including instances of *however*. Thus the rows are candidate wh-formtokens, not a gold-standard set of wh dependencies. The distance is linear andannotation-dependent; it is not a filler-gap distance or a measure of syntacticdepth.:::Let $X_i$ be this selected-token-to-head distance for observation $i$. The observed distances arepositive integers, but the Poisson family also assigns probability to zero. Wewill temporarily ignore that mismatch so that we can isolate the estimationstep. Our working model is$$X_i\mathrel{\overset{\text{IID}}{\sim}}\operatorname{Poisson}(\lambda).$$The model says that every observation shares one Poisson mean $\lambda$ and that theobservations are independent after we condition on that mean. Those claims arestrong. We will check one of their consequences after finding the estimate.## Writing the likelihood one observation at a timeFor one observed distance $x_i$, the Poisson probability mass is$$p(x_i\mid\lambda)=\frac{e^{-\lambda}\lambda^{x_i}}{x_i!}.$$For $n$ independent observations, multiplication gives the joint probability$$p(\mathbf{x}\mid\lambda)=\prod_{i=1}^{n}\frac{e^{-\lambda}\lambda^{x_i}}{x_i!}.$$Once the distances are observed, we hold $\mathbf{x}$ fixed and view the sameexpression as a function of $\lambda$:$$L(\lambda;\mathbf{x})=\prod_{i=1}^{n}\frac{e^{-\lambda}\lambda^{x_i}}{x_i!}.$$This is the likelihood function. It compares candidate values of the mean byasking how much probability each candidate assigns to the observed collectionof distances.## Taking logs before differentiatingThe product contains one small probability for each observation. Taking a logturns that product into a sum:$$\begin{aligned}\ell(\lambda;\mathbf{x})&=\log L(\lambda;\mathbf{x})\\&=\sum_{i=1}^{n}\left[x_i\log\lambda-\lambda-\log(x_i!)\right]\\&=\left(\sum_{i=1}^{n}x_i\right)\log\lambda -n\lambda -\sum_{i=1}^{n}\log(x_i!).\end{aligned}$$The final term depends on the observations but not on $\lambda$. It changes theheight of the log-likelihood curve, but it cannot change where that curve ismaximized.## Finding the maximizing meanDifferentiate with respect to $\lambda$:$$\frac{\mathrm{d}}{\mathrm{d}\lambda}\ell(\lambda;\mathbf{x})=\frac{\sum_{i=1}^{n}x_i}{\lambda}-n.$$At an interior maximum, this derivative is zero. We now isolate $\widehat{\lambda}$ one algebraic move at a time:$$\begin{aligned}\frac{\sum_{i=1}^{n}x_i}{\widehat{\lambda}}-n&=0,\\\sum_{i=1}^{n}x_i&=n\widehat{\lambda},\\\widehat{\lambda}_{\mathrm{MLE}}&=\frac{1}{n}\sum_{i=1}^{n}x_i\\&=\bar{x}.\end{aligned}$$The maximum-likelihood estimate of a shared Poisson mean is the sample mean. Thisequality is a consequence of the Poisson model and its likelihood. It is not ageneral definition of maximum likelihood.## Verifying the result on the dependency distancesThe data are available in[`wh_filler_gap.csv`](../problem-sets/ps1-2/data/wh_filler_gap.csv). We can computethe analytic estimate and compare it with a numerical optimizer.```{r}#| label: poisson-mle-wh-distance#| echo: truewh <-read.csv("../problem-sets/ps1-2/data/wh_filler_gap.csv")x <- wh$linear_distanceanalytic_mle <-mean(x)negative_log_likelihood <-function(lambda, observations) {-sum(dpois(observations, lambda = lambda, log =TRUE))}numerical_fit <-optimize( negative_log_likelihood,interval =c(.01, 20),observations = x)c(analytic = analytic_mle,numerical = numerical_fit$minimum)```Both calculations return approximately $3.007$ tokens. The agreement is not a secondpiece of evidence about the linguistic phenomenon. It is a check that the codeand the algebra solve the same optimization problem.We can also inspect the log-likelihood near the maximum.```{r}#| label: poisson-log-likelihood-curve#| echo: truelambda_grid <-seq(2.5, 3.5, length.out =300)log_likelihood <-vapply( lambda_grid,function(lambda) sum(dpois(x, lambda = lambda, log =TRUE)),numeric(1))plot( lambda_grid, log_likelihood,type ="l",xlab =expression(lambda),ylab ="log-likelihood")abline(v = analytic_mle, lty =2)```The curve is highest at the sample mean. Values near the maximum have similarsupport from the observed data, while more distant values assign lower jointprobability to those observations.## Separating estimation from model adequacyThe estimate answers a conditional question: if the distances were independentdraws from one Poisson distribution, which value of $\lambda$ would fit thembest? It does not establish that the Poisson family is a good description.For this extract,$$\bar{x}\approx3.01\qquad\text{and}\qquads^2\approx10.46.$$A Poisson variable has equal mean and variance. The fitted model must thus haveboth mean and variance near 3.01, while the empirical variance is more thanthree times as large. The distances also cannot be zero in this particularextract, even though the fitted distribution assigns positive probability tozero.These discrepancies identify two distinct model failures: the support mismatchconcerns which values can occur, while the variance mismatch concerns how probabilityis distributed over the values that can occur. Maximum likelihood finds thebest parameter within the chosen family. It cannot repair a family whosestructural assumptions miss the response.::: {.callout-note title="Work through the estimate"}1. For the distances $1,2,2,3,7$, calculate $\widehat{\lambda}_{\mathrm{MLE}}$.2. Explain why $\sum_i\log(x_i!)$ disappears when we differentiate with respect to $\lambda$.3. State the conditional question answered by the estimate.4. Give one reason that selected-token-to-head distances from the same document might not be independent.:::