The preceding page introduced the IID Bernoulli sampling statement. We now use one corpus extract to make that statement concrete before constructing a likelihood. The source is the training portion of the Universal Dependencies English Web Treebank, which contains web text annotated with dependency relations and morphological features.

The extract is retained here because the following maximum-likelihood page uses its exact count and coding. Our immediate question is deliberately narrow: what exactly was counted, and what would one shared Bernoulli probability mean for those rows? This is a continuity example, not a claim that one corpus proportion is a sufficient analysis of pronoun case.

Loading the teaching extract

The stored file contains one row for each personal pronoun token that preprocessing classified as accusative or nonaccusative.

Code
pronoun_case <- read.csv("../slides-archive/lecture7/pronoun_data.csv")

head(pronoun_case)

The response is coded

X_i= \begin{cases} 1 & \text{accusative},\\ 0 & \text{nonaccusative}. \end{cases}

This coding makes the observed sample mean equal to the accusative proportion.

Inspecting the response before summarizing

Code
table(pronoun_case$accusative)

The table contains

n=12{,}280

token observations. Of these,

k=3{,}327

are coded accusative and 8,953 are coded nonaccusative.

The observed proportion is

\widehat{\pi}=\frac{k}{n} =\frac{3327}{12280} \approx.271.

The same calculation in R is

Code
mean(pronoun_case$accusative)

The numerical agreement checks the coding and count calculation.

Stating what one observation represents

One row represents one extracted personal pronoun token. The response records a binary label produced by an older teaching heuristic. Forms that distinguish case on their surface, such as I and me, were labeled from the word form. The syncretic forms you and it were labeled nonaccusative when their basic dependency relation was nsubj and accusative otherwise.

This label is not the UD morphological Case feature. In the English UD guidelines, Case=Nom is called direct case and Case=Acc is called oblique case. By contrast, nsubj and obj name syntactic relations. The fallback treats every non-nsubj occurrence of you or it as accusative, including relations for which that inference is not licensed. We retain the binary file so the likelihood calculation remains reproducible, but refer to the response as the heuristic accusative label throughout.

The row does not represent one speaker, one document, or one pronoun type. Documents and forms repeat across token observations. The web treebank also does not constitute a random sample of all English pronoun uses.

Thus the analysis statement is deliberately narrow:

One observation is one extracted personal pronoun token in the teaching file. The response indicates whether the older preprocessing heuristic assigned that token the accusative label. The immediate estimate is the heuristic-label proportion in this extract.

Declaring the working model

For the calculations that follow, assume

X_i \mathrel{\overset{\mathrm{IID}}{\sim}} \operatorname{Bernoulli}(\pi).

Under this model, \pi is one shared probability of receiving the heuristic accusative label in the represented token sampling process. The independence assumption factors the sample probability, and the identical distribution assumption gives every token the same \pi.

Both claims are strong: pronoun case depends on grammatical function, lexical form, construction, and annotation decisions, while tokens in one document may also be dependent.

The simple model is used to isolate the likelihood calculation. It should not be mistaken for a final linguistic theory.

Separating the response from a theoretical predictor

The binary column records an outcome. It does not explain why one form bears accusative case or why the rate changes across syntactic positions.

A theoretically interesting analysis might compare grammatical functions, local syntactic configurations, pronoun paradigms, or construction types. Those predictors would encode proposals about the distribution of direct and oblique pronoun forms rather than merely reproduce the observed response.

A precise estimate of .271 can describe this extract under the coding rule while leaving the linguistic mechanism unexplained. It is not an estimate of the proportion of tokens bearing Case=Acc in the current UD annotation. The response proportion is a description of the heuristic labels, not an explanation of the case system.

Check your understanding

  1. Verify the observed proportion from k=3{,}327 and n=12{,}280.
  2. What does one row represent, and which higher-level units repeat?
  3. Which two claims are encoded by the IID notation?
  4. Give one theoretically motivated predictor for a substantive pronoun case analysis.

On the maximum-likelihood page, we ask which value of \pi makes this observed sequence most probable under the working model.