Describing what happens

Statistical Methods in Linguistics

Aaron Steven White

University of Rochester

August 31, September 2 and 9, 2026

One utterance, several possible datasets

One production of heed may yield one recording, one word token, one vowel token, or several acoustic measurements.

The analysis begins by deciding which object counts as one observation.

The representation problem

The representation problem is the need to decide what one row represents and which differences among speakers, items, and tokens remain visible.

Linguistic data and comparisons

  1. turn stored records into observations;
  2. turn a research question into a comparison;
  3. distinguish description from process explanation.

Probability

⟨Ω,ℱ,ℙ⟩ \langle\Omega,\mathcal F,\mathbb P\rangle

Ω\Omega represents possible outcomes, ℱ\mathcal F represents measurable events, and ℙ\mathbb P assigns their probabilities.

From observations to probability measures

  1. unit of observation and observation key
  2. linguistic comparison and estimand
  3. possible outcomes
  4. measurable events
  5. probability measure

Each page introduces one object used by the next.

Representation precedes probability

Probability does not decide whether an observation is a file, token, type, or measurement.

That representation is fixed before probabilities are assigned.

More data cannot repair the wrong observation

A numerical result cannot repair a mismatch between the represented observation and the linguistic claim.

More data of the wrong kind remain data of the wrong kind.

Fix the represented observation

Finish the sentence: one observation represents …

Next: identify the unit of observation and observation key.

A recording is not yet an observation

A recording of heed contains a word token, a vowel token, and a sequence of acoustic measurements.

Which one does a row represent?

Unit of observation and observation key

The unit of observation is the entity or event represented by one row.

The decision that fixes this representation is the unit choice.

We reserve unit of analysis for the smallest unit treated as independent or carried into the intended generalization.

We call the smallest collection of columns that uniquely identifies that observation its observation key.

The stored record

A record is a stored artifact: a recording, corpus document, response log, or annotation file.

One record may contain many observations; several records may contribute to one observation.

One reading-time observation

One observation represents one participant’s reading of one region in one sentence.

Its observation key is

⟨participant,sentence,region⟩. \langle\text{participant},\text{sentence},\text{region}\rangle.

What aggregation removes

one row represents variation still visible
participant, sentence, region participants, sentences, regions
participant, sentence participants, sentences
sentence mean sentences

Tokens and types

A token is an occurrence in context. A type groups tokens treated as equivalent.

Ten occurrences of runs are ten tokens of one word-form type; a lemma analysis may group them with run, ran, and running.

A unique row is not an independent row

Two rows may have distinct keys and still share a participant, sentence, speaker, word, or language.

Uniqueness identifies observations; it does not establish independence.

Measurement, analysis, and generalization

  1. Where was the response measured?
  2. What does one analyzed response represent?
  3. Which speakers or items should the claim generalize to?

One thousand tokens from one speaker do not represent variation across speakers.

A pre-analysis specification

One observation represents one attested dative construction token. Its key combines document and token position. Speakers, verbs, and documents repeat.

Specify the unit, key, repeated identifiers, removed variation, and replication before modeling.

Check the observation key

For a token analysis, can two rows share the proposed observation key?

Next: state the comparison that answers the linguistic question.

An illustrative difference is not yet a linguistic result

condition average reading time
noncoercing control 410 ms
coercion condition 455 ms

These teaching values are hypothetical rather than reported results.

What comparison does the 45 ms difference estimate?

455ms−410ms=45ms. 455\ \mathrm{ms}-410\ \mathrm{ms}=45\ \mathrm{ms}.

The comparison-alignment problem

The comparison-alignment problem is the gap between a numerical difference and the claim it is meant to support.

Alignment requires a unit, response, predictor, comparison, and estimand.

Step 1: identify the response

The response is total self-paced reading time, in milliseconds, over the predefined complement-plus-spillover region on one participant-by-sentence trial.

Step 2: identify the linguistic contrast

The main predictor distinguishes coercion from noncoercing control trials.

If the conditions use different verb sets, it does not isolate coercion.

Step 3: identify the compared observations

Compare conditions within participants and within sentence items.

Different participant groups or item sets introduce additional differences.

Step 4: state the estimand

Let μcoercion\mu_{\mathrm{coercion}} and μcontrol\mu_{\mathrm{control}} be target-population mean reading times.

τ≡μcoercion−μcontrol. \tau \equiv \mu_{\mathrm{coercion}} - \mu_{\mathrm{control}}.

The complete analysis specification

One observation is one participant-by-sentence reading of the declared region; the comparison is within participants and items; the target is τ\tau.

The same records can answer different questions

The same trials may support analyses of reading time, interpretation choice, or associations between participant-level contrasts.

These are different estimands.

Three alignment failures

  1. One speaker’s tokens cannot estimate variation among speakers.
  2. Different verb sets do not isolate coercion.
  3. Trials from observed readers do not test transfer to new readers.

Changing the estimator repairs none of these failures.

Exploration can change the question

If slider responses pile up at the endpoints, a target concerned only with mean position may omit the endpoint-versus-interior distinction.

Revise and report the estimand explicitly.

Check the comparison alignment

Does the response, predictor, comparison, and estimand refer to the same unit and population?

Next: apply the sequence to word frequency and rank.

Count occurrences or groups?

Washington’s 1793 address contains 135 represented word tokens assigned to 90 orthographic word-form types.

Frequency counts tokens; rank orders types.

Zipf’s law

Zipf’s law names an approximate inverse relation between a represented type’s frequency and its frequency rank.

Observed value, prediction, residual

f(r)≡token count at rank r, f(r)\equiv\text{token count at rank }r,

f̂(r)≡Cr−α,C>0,α>0,e(r)≡f(r)−f̂(r). \widehat f(r)\equiv Cr^{-\alpha}, \qquad C>0,\ \alpha>0, \qquad e(r)\equiv f(r)-\widehat f(r).

Count the tokens

type the of I to in shall
count 13 11 6 5 3 3

Ties are ordered alphabetically so that the ranks are reproducible.

The counting representation

State what counts as a token, a type, and the corpus; then state the tie rule.

Lemmatizing run, runs, ran, and running changes counts and ranks.

Compare counts with the inverse pattern

Fix α=1\alpha=1 and anchor the curve at rank 1, so C=f(1)=13C=f(1)=13.

rr f(r)f(r) f̂(r)=13/r\widehat f(r)=13/r e(r)e(r)
1 13 13.00 0.00
2 11 6.50 4.50
3 6 4.33 1.67

We do not estimate α\alpha or compare explanations of the rank-frequency relation.

The represented counts and inverse prediction

Word frequency by frequency rank for Washington's 1793 inaugural address, with the inverse-rank prediction shown as a line.

The points are observed type counts; the line is 13/r13/r.

Inspect the broad pattern

Normalized word frequency plotted against frequency rank.

A fitted curve describes the rank-frequency relation in the represented corpus.

Piantadosi 2014, Figure 1a

Inspect the departures from the curve

Observed deviations from a simple Zipfian rank-frequency curve.

The departures are systematic, especially among very frequent and very infrequent types.

Piantadosi 2014, Figure 1b

The description-explanation gap

Does an inverse rank-frequency curve describe the broad pattern?

What process produced the particular frequencies and deviations in this corpus?

The descriptive target

One observation represents one word-form type; its response is token count; rank follows from sorting under the declared tie rule.

A fitted curve does not explain lexical choice, memory, or communication.

Declare tokens, types, corpus, and rank ties

Which choices define the tokens, types, corpus, and tie ordering?

Next: declare the possible outcomes.

Which result does the model represent?

One pronoun token may be represented by its recording, surface form, or annotated form.

The probability model must choose one.

The outcome-representation problem

The outcome-representation problem is the need to decide which possible results are members of the model’s sample space.

Outcome and sample space

An outcome is one possible represented result. The sample space is the set of all outcomes:

Ω≡{𝐼,𝑚𝑒,𝑦𝑜𝑢,𝑖𝑡,ℎ𝑒,ℎ𝑖𝑚,𝑠ℎ𝑒,ℎ𝑒𝑟,𝑤𝑒,𝑢𝑠,𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚}. \Omega\equiv \{\textit{I},\textit{me},\textit{you},\textit{it},\textit{he},\textit{him}, \textit{she},\textit{her},\textit{we},\textit{us},\textit{they},\textit{them}\}.

Membership and a trial record

ℎ𝑖𝑚∈Ω,ℎ𝑒𝑟𝑠𝑒𝑙𝑓∉Ω. \textit{him}\in\Omega, \qquad \textit{herself}\notin\Omega.

A response-retaining outcome may instead be

⟨participant and sentence trial,6⟩. \langle\text{participant and sentence trial},6\rangle.

Outcome granularity fixes the available distinctions

Outcome granularity determines which distinctions remain available: recording, vowel token, time point, response value, or identified trial.

Omission is not impossibility

If skipping is recordable, then

{agent,patient} \{\text{agent},\text{patient}\}

omits a possible result. The omission does not make skipping impossible.

Check the modeled outcomes

Does ℎ𝑒𝑟𝑠𝑒𝑙𝑓∉Ω\textit{herself}\notin\Omega describe English, or only this model?

Next: group outcomes into events.

Linguistic questions group outcomes

Was the selected pronoun third person? Was it nominative? Was it accusative?

Each yes-or-no question collects the outcomes that make it true.

Event

An event is a subset of Ω\Omega. Once ℱ\mathcal F is specified, its members are the measurable events.

Person and case in the 12-form space

T12≡{𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚,𝑖𝑡,𝑠ℎ𝑒,ℎ𝑒𝑟,ℎ𝑒,ℎ𝑖𝑚}, T_{12}\equiv\{\textit{they},\textit{them},\textit{it},\textit{she}, \textit{her},\textit{he},\textit{him}\},

A12≡{𝑚𝑒,𝑦𝑜𝑢,𝑡ℎ𝑒𝑚,ℎ𝑒𝑟,ℎ𝑖𝑚,𝑖𝑡,𝑢𝑠},N12≡{𝐼,𝑦𝑜𝑢,𝑡ℎ𝑒𝑦,𝑠ℎ𝑒,ℎ𝑒,𝑖𝑡,𝑤𝑒}. \begin{aligned} A_{12}&\equiv\{\textit{me},\textit{you},\textit{them},\textit{her}, \textit{him},\textit{it},\textit{us}\},\\ N_{12}&\equiv\{\textit{I},\textit{you},\textit{they},\textit{she}, \textit{he},\textit{it},\textit{we}\}. \end{aligned}

A12∩N12={𝑦𝑜𝑢,𝑖𝑡}. A_{12}\cap N_{12}=\{\textit{you},\textit{it}\}.

Surface form alone does not resolve the case of you or it.

A finite number-and-case space

Ω4≡{ℎ𝑒,ℎ𝑖𝑚,𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚}. \Omega_4 \equiv\{\textit{he},\textit{him},\textit{they},\textit{them}\}.

One outcome is one case-coded pronoun token from this stipulated four-form population.

Defining the number and case events

P4≡{𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚},A4≡{ℎ𝑖𝑚,𝑡ℎ𝑒𝑚}. P_4\equiv\{\textit{they},\textit{them}\}, \qquad A_4\equiv\{\textit{him},\textit{them}\}.

Intersection

P4∩A4={𝑡ℎ𝑒𝑚}. P_4\cap A_4=\{\textit{them}\}.

The selected token is both plural and accusative.

Union

P4∪A4={ℎ𝑖𝑚,𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚}. P_4\cup A_4=\{\textit{him},\textit{they},\textit{them}\}.

The union is inclusive: them belongs because it is in both events.

Complement

P4c=Ω4\P4={ℎ𝑒,ℎ𝑖𝑚}. P_4^c=\Omega_4\setminus P_4=\{\textit{he},\textit{him}\}.

Complements are relative to the declared sample space.

Event membership

One outcome can belong to several events. The outcome them makes both P4P_4 and A4A_4 occur.

The limiting events

⌀\varnothing contains no outcomes; Ω4\Omega_4 contains every represented outcome.

Both belong to every sigma-algebra on Ω4\Omega_4.

Co-occurrence depends on the outcome

Nominative and accusative are disjoint when one outcome is one case-coded token.

They may co-occur when one outcome is a sentence containing several pronouns.

Compute the union before assigning probability

Compute P4∪A4P_4\cup A_4 before assigning any probabilities.

Next: decide which events belong to the sigma-algebra.

Which questions must remain measurable?

If plurality is measurable, its complement must be measurable. If two events are measurable, their union must be measurable.

This is the closure problem.

Sigma-algebra

A sigma-algebra ℱ\mathcal F contains Ω\Omega and is closed under complements and countable unions.

The person event space

ℱperson≡σ(T12)={⌀,T12,T12c,Ω}. \mathcal F_{\mathrm{person}} \equiv\sigma(T_{12}) =\{\varnothing,T_{12},T_{12}^c,\Omega\}.

The case event space

The three atoms are

A12∩N12,A12\N12,N12\A12. A_{12}\cap N_{12}, \qquad A_{12}\setminus N_{12}, \qquad N_{12}\setminus A_{12}.

Their eight unions form σ(A12,N12)\sigma(A_{12},N_{12}).

ℱcase≡σ(A12,N12). \mathcal F_{\mathrm{case}} \equiv\sigma(A_{12},N_{12}).

For P4={𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚}P_4=\{\textit{they},\textit{them}\},

ℱP4≡{⌀,P4,P4c,Ω4}. \mathcal F_{P_4} \equiv \{\varnothing,P_4,P_4^c,\Omega_4\}.

Closure requirements

  1. Ω∈ℱ\Omega\in\mathcal F.
  2. If B∈ℱB\in\mathcal F, then Bc∈ℱB^c\in\mathcal F.
  3. If B1,B2,…∈ℱB_1,B_2,\ldots\in\mathcal F, then ⋃iBi∈ℱ\bigcup_iB_i\in\mathcal F.

De Morgan’s law also licenses intersections:

B∩C=(Bc∪Cc)c. B\cap C=(B^c\cup C^c)^c.

A collection that is not a sigma-algebra

𝒢≡{Ω4,P4} \mathcal G\equiv\{\Omega_4,P_4\}

is not a sigma-algebra: P4c∉𝒢P_4^c\notin\mathcal G and ⌀∉𝒢\varnothing\notin\mathcal G.

Sample spaces and event spaces

Ω\Omega lists represented outcomes. ℱ\mathcal F lists the outcome groupings that may receive probabilities.

These are different modeling choices.

An event outside the sigma-algebra

A4∉ℱP4A_4\notin\mathcal F_{P_4}. A probability measure with domain ℱP4\mathcal F_{P_4} thus cannot assign ℙ(A4)\mathbb P(A_4).

Subset does not imply measurability

Is {𝑡ℎ𝑒𝑚}⊆Ω4\{\textit{them}\}\subseteq\Omega_4 enough to guarantee {𝑡ℎ𝑒𝑚}∈ℱP4\{\textit{them}\}\in\mathcal F_{P_4}?

Next: generate the smallest coherent event space.

Two event spaces do not combine by union

T12∈ℱpersonT_{12}\in\mathcal F_{\mathrm{person}} and A12∈ℱcaseA_{12}\in\mathcal F_{\mathrm{case}}, but

T12∩A12∉ℱperson∪ℱcase. T_{12}\cap A_{12} \notin \mathcal F_{\mathrm{person}}\cup\mathcal F_{\mathrm{case}}.

The generation problem

The generation problem asks for the smallest sigma-algebra containing chosen event distinctions:

ℱperson-case≡σ(T12,A12,N12). \mathcal F_{\mathrm{person\text{-}case}} \equiv\sigma(T_{12},A_{12},N_{12}).

{T12,A12,N12}\{T_{12},A_{12},N_{12}\} is the generating set.

Refine syncretic outcomes

Ω14≡{𝐼,𝑚𝑒,𝑦𝑜𝑢[−acc],𝑦𝑜𝑢[+acc],𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚,𝑖𝑡[−acc],𝑖𝑡[+acc],𝑠ℎ𝑒,ℎ𝑒𝑟,ℎ𝑒,ℎ𝑖𝑚,𝑤𝑒,𝑢𝑠}. \begin{aligned} \Omega_{14}\equiv\{&\textit{I},\textit{me}, \textit{you}_{[-\mathrm{acc}]},\textit{you}_{[+\mathrm{acc}]}, \textit{they},\textit{them},\\ &\textit{it}_{[-\mathrm{acc}]},\textit{it}_{[+\mathrm{acc}]}, \textit{she},\textit{her},\textit{he},\textit{him},\textit{we},\textit{us}\}. \end{aligned}

Accusative outcomes in the refined space

A14≡{𝑚𝑒,𝑦𝑜𝑢[+acc],𝑡ℎ𝑒𝑚,ℎ𝑒𝑟,ℎ𝑖𝑚,𝑖𝑡[+acc],𝑢𝑠}. A_{14}\equiv\{\textit{me},\textit{you}_{[+\mathrm{acc}]},\textit{them}, \textit{her},\textit{him},\textit{it}_{[+\mathrm{acc}]},\textit{us}\}.

Third-person outcomes in the refined space

T14≡{𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚,𝑖𝑡[−acc],𝑖𝑡[+acc],𝑠ℎ𝑒,ℎ𝑒𝑟,ℎ𝑒,ℎ𝑖𝑚}. T_{14}\equiv\{\textit{they},\textit{them}, \textit{it}_{[-\mathrm{acc}]},\textit{it}_{[+\mathrm{acc}]}, \textit{she},\textit{her},\textit{he},\textit{him}\}.

Four atoms from two generators

T14∩A14,T14∩A14c,T14c∩A14,T14c∩A14c. T_{14}\cap A_{14},\quad T_{14}\cap A_{14}^c,\quad T_{14}^c\cap A_{14},\quad T_{14}^c\cap A_{14}^c.

Their 24=162^4=16 unions form σ(T14,A14)\sigma(T_{14},A_{14}).

A finite pronoun space

Ω4≡{ℎ𝑒,ℎ𝑖𝑚,𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚}, \Omega_4\equiv\{\textit{he},\textit{him},\textit{they},\textit{them}\},

P4≡{𝑡ℎ𝑒𝑦,𝑡ℎ𝑒𝑚},A4≡{ℎ𝑖𝑚,𝑡ℎ𝑒𝑚}. P_4\equiv\{\textit{they},\textit{them}\}, \qquad A_4\equiv\{\textit{him},\textit{them}\}.

Atoms of the generated sigma-algebra

A4A4cP4{𝑡ℎ𝑒𝑚}{𝑡ℎ𝑒𝑦}P4c{ℎ𝑖𝑚}{ℎ𝑒} \begin{array}{c|cc} &A_4&A_4^c\\\hline P_4&\{\textit{them}\}&\{\textit{they}\}\\ P_4^c&\{\textit{him}\}&\{\textit{he}\} \end{array}

Thus σ(P4,A4)=2Ω4\sigma(P_4,A_4)=2^{\Omega_4}.

Inspect the finite construction

{P4,A4}⇒4 singleton atoms⇒24=16 unions. \{P_4,A_4\} \quad\Longrightarrow\quad 4\text{ singleton atoms} \quad\Longrightarrow\quad 2^4=16\text{ unions}.

Coarse generators preserve coarse atoms

σ(P4)={⌀,P4,P4c,Ω4}. \sigma(P_4)=\{\varnothing,P_4,P_4^c,\Omega_4\}.

This event space cannot separate they from them and cannot represent case.

When generators omit a needed distinction

If two outcomes have the same membership in every generator, no generated event can separate them.

More tokens cannot add case to σ(P4)\sigma(P_4); the model needs a case generator.

A finite construction procedure

Choose each generator or its complement; intersect; remove empty cells; take every union of the remaining atoms.

Check the generators and atoms

Can two outcomes with identical generator membership be separated?

Next: assign probability measures to the measurable events.

Measurable events still need weights

Ω\Omega says what may happen, and ℱ\mathcal F says which groupings are measurable.

How can numbers be assigned without contradiction?

Probability measure and probability space

A probability measure is a function

ℙ:ℱ→[0,1]. \mathbb P:\mathcal F\to[0,1].

The triple ⟨Ω,ℱ,ℙ⟩\langle\Omega,\mathcal F,\mathbb P\rangle is a probability space.

A uniform Ω14\Omega_{14} calculation

For E∈σ(T14,A14)E\in\sigma(T_{14},A_{14}), let ℙ(E)=|E|/14\mathbb P(E)=|E|/14. Then

ℙ(T14∩A14)=414=27. \mathbb P(T_{14}\cap A_{14}) =\frac{4}{14} =\frac27.

Event space and measure do different work

ℱ\mathcal F determines which questions receive probabilities. ℙ\mathbb P supplies the numerical answers.

Neither one determines the other.

A nonuniform Ω4\Omega_4 calculation

outcome he him they them
mass .20.20 .30.30 .10.10 .40.40

ℙ(A4)=.30+.40=.70. \mathbb P(A_4)=.30+.40=.70.

The coherence problem

The coherence problem is to keep probabilities nonnegative, normalized, and additive. For pairwise-disjoint Bi∈ℱB_i\in\mathcal F,

ℙ(B)≥0,ℙ(Ω)=1,ℙ(⋃iBi)=∑iℙ(Bi). \mathbb P(B)\geq0, \quad \mathbb P(\Omega)=1, \quad \mathbb P\!\left(\bigcup_iB_i\right)=\sum_i\mathbb P(B_i).

Verify the toy measure

.20+.30+.10+.40=1,ℙ({ℎ𝑒,ℎ𝑖𝑚})=.50. .20+.30+.10+.40=1, \qquad \mathbb P(\{\textit{he},\textit{him}\})=.50.

Nonnegativity, normalization, and additivity all hold.

Two consequences of the axioms

ℙ(⌀)=0,ℙ(Bc)=1−ℙ(B). \mathbb P(\varnothing)=0, \qquad \mathbb P(B^c)=1-\mathbb P(B).

An assignment that violates the axioms

Masses .20,.30,.10,.50.20,.30,.10,.50 sum to 1.101.10 and thus violate normalization.

Check the probability measure

Do the assigned masses satisfy nonnegativity, normalization, and additivity?

Next: determine when event probabilities may be added without overlap.

Can two event probabilities be added?

In Ω4\Omega_4, nominative and accusative forms cannot be the same selected token.

Plural and accusative forms can.

Mutual exclusivity

Events BB and CC are mutually exclusive exactly when

B∩C=⌀. B\cap C=\varnothing.

Addition rules

If B∩C=⌀B\cap C=\varnothing,

ℙ(B∪C)=ℙ(B)+ℙ(C). \mathbb P(B\cup C)=\mathbb P(B)+\mathbb P(C).

In general,

ℙ(B∪C)=ℙ(B)+ℙ(C)−ℙ(B∩C). \mathbb P(B\cup C)=\mathbb P(B)+\mathbb P(C)-\mathbb P(B\cap C).

A disjoint calculation

N4={ℎ𝑒,𝑡ℎ𝑒𝑦},A4={ℎ𝑖𝑚,𝑡ℎ𝑒𝑚}. N_4=\{\textit{he},\textit{they}\}, \qquad A_4=\{\textit{him},\textit{them}\}.

ℙ(N4∪A4)=.30+.70=1. \mathbb P(N_4\cup A_4)=.30+.70=1.

Disjointness and exhaustivity differ

Disjointness licenses addition. Exhaustivity explains why a disjoint sum equals one.

The two properties are distinct.

The double-counting problem

P4∩A4={𝑡ℎ𝑒𝑚}P_4\cap A_4=\{\textit{them}\}, so

ℙ(P4)+ℙ(A4)=.50+.70=1.20 \mathbb P(P_4)+\mathbb P(A_4)=.50+.70=1.20

counts the mass on them twice.

Dependence on the represented outcome

Nominative and accusative are disjoint when one outcome is one token.

The corresponding sentence-level events may co-occur when a sentence contains two pronouns.

Labels do not guarantee mutual exclusivity

Contrasting linguistic labels do not establish an empty intersection.

Write the events as sets and inspect their overlap.

Inspect overlap before adding probabilities

Write the event sets and inspect their intersection before adding probabilities.

Next: assign probability directly to an intersection.

How likely is both plural and accusative?

The question concerns outcomes shared by the plural and accusative events.

That shared event is an intersection.

The original uniform example

T14∩A14={𝑡ℎ𝑒𝑚,𝑖𝑡[+acc],ℎ𝑒𝑟,ℎ𝑖𝑚}, T_{14}\cap A_{14} =\{\textit{them},\textit{it}_{[+\mathrm{acc}]},\textit{her},\textit{him}\},

ℙ(T14,A14)≡ℙ(T14∩A14)=4/14=2/7. \mathbb P(T_{14},A_{14}) \equiv\mathbb P(T_{14}\cap A_{14}) =4/14=2/7.

Joint probability

The joint probability of BB and CC is defined by

ℙ(B,C)≡ℙ(B∩C). \mathbb P(B,C) \equiv \mathbb P(B\cap C).

Full-space calculation

P4∩A4={𝑡ℎ𝑒𝑚} P_4\cap A_4=\{\textit{them}\}

ℙ(P4,A4)=ℙ({𝑡ℎ𝑒𝑚})=.40. \mathbb P(P_4,A_4) =\mathbb P(\{\textit{them}\}) =.40.

The joint table

A4A_4 A4cA_4^c total
P4P_4 .40.40 .10.10 .50.50
P4cP_4^c .30.30 .20.20 .50.50
total .70.70 .30.30 11

Reconstruct the table

row sums=(.50,.50),column sums=(.70,.30),total=1. \text{row sums}=(.50,.50), \qquad \text{column sums}=(.70,.30), \qquad \text{total}=1.

Marginal probabilities

Marginalization sums over exhaustive case alternatives. The row and column totals are marginal probabilities:

ℙ(P4)=ℙ(P4,A4)+ℙ(P4,A4c)=.50. \mathbb P(P_4) =\mathbb P(P_4,A_4)+\mathbb P(P_4,A_4^c) =.50.

Information in the joint table

Two measures may share ℙ(P4)=.50\mathbb P(P_4)=.50 and ℙ(A4)=.70\mathbb P(A_4)=.70 while assigning different mass to P4∩A4P_4\cap A_4.

Marginal and joint probabilities

ℙ(A4)=.70\mathbb P(A_4)=.70 includes him and them.

Only the .40.40 assigned to them belongs to P4∩A4P_4\cap A_4.

Distinguish joint from conditional probability

Why does ℙ(P4,A4)=.40\mathbb P(P_4,A_4)=.40 not mean that 40% of plural forms are accusative?

Next: change the reference event with conditional probability.

Among plural forms, how many are accusative?

The joint probability uses the full sample space.

The word among changes the reference event to P4P_4.

Conditional probability

For B,C∈ℱB,C\in\mathcal F with ℙ(C)>0\mathbb P(C)>0,

ℙ(B∣C)≡ℙ(B,C)ℙ(C). \mathbb P(B\mid C) \equiv \frac{\mathbb P(B,C)}{\mathbb P(C)}.

A uniform Ω14\Omega_{14} calculation

ℙ(T14∣A14)=4/147/14=47. \mathbb P(T_{14}\mid A_{14}) =\frac{4/14}{7/14} =\frac47.

The denominator is the accusative reference event.

Rescale the plural event

outcome mass in Ω4\Omega_4 mass within P4P_4
they .10.10 .10/.50=.20.10/.50=.20
them .40.40 .40/.50=.80.40/.50=.80

ℙ(A4∣P4)=.80. \mathbb P(A_4\mid P_4)=.80.

Conditioning as rescaling

Conditioning retains the mass inside CC and rescales it to one.

ℙ(C∣C)=1. \mathbb P(C\mid C)=1.

Reversing target and condition

ℙ(P4∣A4)=.40.70≈.571≠ℙ(A4∣P4). \mathbb P(P_4\mid A_4)=\frac{.40}{.70}\approx.571 \neq \mathbb P(A_4\mid P_4).

The numerator is unchanged; the denominator changes.

Positive-probability conditions

If ℙ(C)=0\mathbb P(C)=0, the elementary ratio is undefined.

The condition ℙ(C)>0\mathbb P(C)>0 licenses division by the reference-event probability.

The order of the events

ℙ(animate subject∣passive)\mathbb P(\text{animate subject}\mid\text{passive}) does not determine ℙ(passive∣animate subject)\mathbb P(\text{passive}\mid\text{animate subject}).

Read the event to the right of the bar first.

Name the conditioning event

State the denominator of ℙ(animate subject∣passive)\mathbb P(\text{animate subject}\mid\text{passive}) in words.

Next: rearrange the definition to obtain the chain rule.

Recover the joint probability

Half of the represented tokens are plural; 80% of the plural tokens are accusative.

What proportion are both?

Factorization and the multiplication rule

A factorization rewrites a joint probability as a product. The multiplication rule, also called the two-event chain rule, is

ℙ(B,C)=ℙ(B∣C)ℙ(C). \mathbb P(B,C) =\mathbb P(B\mid C)\mathbb P(C).

Derive the two-event rule

ℙ(B∣C)=ℙ(B,C)ℙ(C) \mathbb P(B\mid C) =\frac{\mathbb P(B,C)}{\mathbb P(C)}

⇒ℙ(B,C)=ℙ(B∣C)ℙ(C). \Longrightarrow \qquad \mathbb P(B,C) =\mathbb P(B\mid C)\mathbb P(C).

Reconstructing the joint probability

ℙ(A4,P4)=.80×.50=.40. \mathbb P(A_4,P_4) =.80\times.50 =.40.

Among 100 tokens: 50 are plural, and 40 of those 50 are accusative.

Reversing the order

ℙ(A4,P4)=ℙ(P4∣A4)ℙ(A4)=.40.70(.70)=.40. \mathbb P(A_4,P_4) =\mathbb P(P_4\mid A_4)\mathbb P(A_4) =\frac{.40}{.70}(.70)=.40.

Three-event chain rule

ℙ(B,C,D)=ℙ(B)ℙ(C∣B)ℙ(D∣B,C). \mathbb P(B,C,D) =\mathbb P(B)\mathbb P(C\mid B)\mathbb P(D\mid B,C).

The two-event rule is applied twice.

General chain rule

When every displayed conditioning event has positive probability,

ℙ(E1,…,EN)=ℙ(E1)∏i=2Nℙ(Ei∣E1,…,Ei−1). \mathbb P(E_1,\ldots,E_N) =\mathbb P(E_1) \prod_{i=2}^{N} \mathbb P(E_i\mid E_1,\ldots,E_{i-1}).

Factorization and dependence

The conditional factors retain dependence. Replacing ℙ(B∣C)\mathbb P(B\mid C) with ℙ(B)\mathbb P(B) requires independence.

The chain rule itself does not.

A linguistic application

ℙ(short duration,high predictability)=ℙ(short duration∣high predictability)ℙ(high predictability). \begin{aligned} &\mathbb P(\text{short duration},\text{high predictability})\\ &\quad=\mathbb P(\text{short duration}\mid\text{high predictability}) \mathbb P(\text{high predictability}). \end{aligned}

This factorization represents an association in the chosen population. It does not show that predictability caused shorter duration.

Expand the four-event chain rule

Write the chain rule for ℙ(E1,E2,E3,E4)\mathbb P(E_1,E_2,E_3,E_4) in that order.

Next: factor one joint probability twice to derive Bayes’ rule.

Reverse a conditional direction

A pronunciation feature may be common within a small dialect group.

Let DD be dialect-group membership and FF be use of the pronunciation feature.

Does that imply that most tokens containing the feature come from that group?

The conditional-direction problem

The conditional-direction problem is that ℙ(F∣D)\mathbb P(F\mid D) does not determine ℙ(D∣F)\mathbb P(D\mid F) by itself.

Prevalence and overall feature probability also matter.

Factor the same joint probability twice

ℙ(B,C)=ℙ(B∣C)ℙ(C)=ℙ(C∣B)ℙ(B). \mathbb P(B,C)=\mathbb P(B\mid C)\mathbb P(C) =\mathbb P(C\mid B)\mathbb P(B).

Bayes’ rule

Factoring the same joint probability twice gives

ℙ(B∣C)=ℙ(C∣B)ℙ(B)ℙ(C). \mathbb P(B\mid C) = \frac{\mathbb P(C\mid B)\mathbb P(B)} {\mathbb P(C)}.

The pronoun measure

ℙ(A4∣P4)=(.40/.70)(.70).50=.80. \mathbb P(A_4\mid P_4) =\frac{(.40/.70)(.70)}{.50} =.80.

The reversed conditional and the two margins recover the original direction.

A hypothetical dialect-feature table

feature FF no feature FcF^c total
dialect group DD 90 10 100
other group DcD^c 180 720 900
total 270 730 1,000

ℙ(D∣F)=.90(.10).27=13. \mathbb P(D\mid F)=\frac{.90(.10)}{.27}=\frac13.

Evidence, prevalence, and sampling

ℙ(F∣D)\mathbb P(F\mid D) describes the feature within the group; ℙ(D)\mathbb P(D) is group prevalence; ℙ(F)\mathbb P(F) is overall feature probability.

A balanced sample does not directly provide population prevalence.

The order of the events

ℙ(F∣D)=90/100\mathbb P(F\mid D)=90/100 uses dialect-group tokens as its denominator.

ℙ(D∣F)=90/270\mathbb P(D\mid F)=90/270 uses feature tokens as its denominator.

Reverse the conditional

Why do .90.90 and 1/31/3 answer different questions?

Next: ask when conditioning leaves a probability unchanged in independence.

Does conditioning change the probability?

ℙ(A4∣P4)=.80butℙ(A4)=.70. \mathbb P(A_4\mid P_4)=.80 \qquad\text{but}\qquad \mathbb P(A_4)=.70.

Plurality changes the probability of accusative form.

The invariance test

The invariance test compares a probability before and after conditioning:

ℙ(B∣C)=?ℙ(B). \mathbb P(B\mid C) \stackrel{?}{=} \mathbb P(B).

Independence

Events BB and CC are independent exactly when

B⟂⟂C≡[ℙ(B,C)=ℙ(B)ℙ(C)]. B\perp\!\!\!\perp C \quad\equiv\quad \bigl[\mathbb P(B,C)=\mathbb P(B)\mathbb P(C)\bigr].

When ℙ(C)>0\mathbb P(C)>0, this is equivalent to conditional invariance.

Independence in the original uniform example

ℙ(T14∣A14)=47=ℙ(T14), \mathbb P(T_{14}\mid A_{14})=\frac47=\mathbb P(T_{14}),

ℙ(T14,A14)=414=814714. \mathbb P(T_{14},A_{14}) =\frac4{14} =\frac8{14}\frac7{14}.

The pronoun measure

ℙ(P4,A4)=.40≠.50(.70)=.35. \mathbb P(P_4,A_4)=.40 \neq .50(.70)=.35.

Plurality and accusative case are dependent under this measure.

Equivalence with conditional invariance

For ℙ(C)>0\mathbb P(C)>0,

ℙ(B∣C)=ℙ(B)⇔ℙ(B,C)=ℙ(B)ℙ(C). \mathbb P(B\mid C)=\mathbb P(B) \quad\Longleftrightarrow\quad \mathbb P(B,C)=\mathbb P(B)\mathbb P(C).

An independent comparison measure

measure ℙ(P4,A4)\mathbb P(P_4,A_4) ℙ(P4)ℙ(A4)\mathbb P(P_4)\mathbb P(A_4) independent?
comparison .35.35 .35.35 yes

Here ℙ(A4∣P4)=.35/.50=.70=ℙ(A4)\mathbb P(A_4\mid P_4)=.35/.50=.70=\mathbb P(A_4).

The same events under different measures

Independence is a property of events together with a probability measure.

The same linguistic labels may be independent in one population and dependent in another.

Independence is not mutual exclusivity

Mutually exclusive positive-probability events are not independent: observing one rules out the other.

Sample estimates and independence

Finite-sample proportions rarely satisfy the independence equation exactly, even under population independence.

They may also agree by chance when the events are dependent.

Check and module upshot

Can a finite sample miss a dependence or suggest one by chance? Yes.

Represent first; then define events; then calculate under the declared measure.