From outcomes to distributions

Statistical Methods in Linguistics

Aaron Steven White

University of Rochester

September 14, 16, 21 and 23, 2026

One outcome can answer several questions

A contextualized pronoun token has a spelling, an annotated case value, a speaker, and a duration. Which part of that record should enter a particular analysis?

The outcome-to-value question

Module 1 assigned probability to events, that is, observable sets of complete outcomes.

But a linguistic question usually targets one attribute of each outcome:

  • its orthographic form;
  • its annotated morphosyntactic case; or
  • a measured duration.

Preserve the relevant events without discarding the outcome

If an analysis concerns duration, then replacing each complete token by its duration should preserve the events needed for duration questions.

But it should not erase the token record that supports other questions.

The outcome-to-value mapping

A random variable extracts one value from each outcome. Its preimages connect value questions back to the event space from Module 1.

Review the notes: why random variables? · Module 1: events

From outcomes to joint distributions

  1. September 14: Random variables, PMFs, densities, and CDFs
  2. September 16: Expectations, variance, and moments
  3. September 21: Standard distribution families
  4. September 23: Joint, marginal, and conditional distributions

Each meeting starts from a linguistic question, defines the probability object that answers it, and checks the definition on a linguistic example.

Random variables and distributions

September 14

Complete vowel-token outcomes

Suppose the random experiment selects one complete vowel token.

outcome word speaker duration
ω1\omega_1 heed s01 112 ms
ω2\omega_2 hid s01 83 ms
ω3\omega_3 heed s02 126 ms

The outcome retains the word, speaker, waveform, vowel label, and duration.

Module 1: possibilities and outcomes

The duration mapping

Let Ω≡{ω1,ω2,ω3}\Omega\equiv\{\omega_1,\omega_2,\omega_3\}. Define

D(ω)≡duration of vowel token ω in milliseconds. D(\omega)\equiv\text{duration of vowel token }\omega\text{ in milliseconds}.

Thus D(ω1)=112D(\omega_1)=112, D(ω2)=83D(\omega_2)=83, and D(ω3)=126D(\omega_3)=126.

Duration values define events

The event that duration exceeds 100 ms is

{ω∈Ω∣D(ω)>100}={ω1,ω3}. \{\omega\in\Omega\mid D(\omega)>100\} =\{\omega_1,\omega_3\}.

The event contains vowel-token outcomes, not the bare values 112 and 126.

One outcome supports two mappings

Define

D(ω)≡duration in ms,H(ω)≡{1,if the vowel is high,0,otherwise. D(\omega)\equiv\text{duration in ms}, \qquad H(\omega)\equiv \begin{cases} 1,&\text{if the vowel is high},\\ 0,&\text{otherwise}. \end{cases}

The same outcome may satisfy D(ω1)=112D(\omega_1)=112 and H(ω1)=1H(\omega_1)=1.

Checking the selection in R

tokens <- data.frame(
  outcome = c("omega1", "omega2", "omega3"),
  word = c("heed", "hid", "heed"),
  speaker = c("s01", "s01", "s02"),
  duration = c(112, 83, 126)
)
tokens$outcome[tokens$duration > 100]
# "omega1" "omega3"

The operation selects duration values but returns the corresponding outcome identifiers.

Retaining the complete outcomes

Keeping only 112, 83, and 126 removes speaker and word identity.

A later analysis could no longer ask whether durations repeat within speakers or differ across words.

Outcome-to-value upshot

A random variable maps complete outcomes to values. Retaining the outcome identifiers keeps each extracted value connected to the observation process.

Notes: why random variables?

Defining a random variable

Basically, a random variable is a rule that reads one value from each outcome.

Write the outcome-to-value mapping as

X:Ω→𝒳. X:\Omega\rightarrow\mathcal X.

The domain Ω\Omega contains outcomes. The codomain 𝒳\mathcal X contains possible values. We add the formal measurability requirement after examining preimages.

Full definition in the notes · OpenStax: random variables

The declared codomain

For a random variable XX, write its declared value space as

cod⁡(X). \operatorname{cod}(X).

Thus a sum over possible values uses x∈cod⁡(X)x\in\operatorname{cod}(X). The support may be smaller.

For a discrete variable, its support is the subset of codomain values with positive probability mass.

Work one mapping by hand

For each token ω∈Ω\omega\in\Omega, define

D(ω)≡duration of token ω in milliseconds. D(\omega)\equiv\text{duration of token }\omega\text{ in milliseconds}.

Thus D(ω1)=112D(\omega_1)=112, D(ω2)=83D(\omega_2)=83, and D(ω3)=126D(\omega_3)=126.

The numbers are measured values in milliseconds. They are not the speech tokens themselves.

How does a value question become an event?

Suppose we ask whether a selected token lasts more than 100 ms.

Prediction: the answer should be a set of token outcomes, since Module 1 assigns probability to events rather than to ungrounded values.

Value statements denote preimage events

For a set of values B⊆𝒳B\subseteq\mathcal X, define the shorthand

X−1(B)={ω∈Ω∣X(ω)∈B}. X^{-1}(B) =\{\omega\in\Omega\mid X(\omega)\in B\}.

This set contains outcomes, not bare values.

The preimage X−1(B)X^{-1}(B) gathers exactly the outcomes whose values land in BB.

A threshold event

For the three-outcome illustration, let Ω≡{ω1,ω2,ω3}\Omega\equiv\{\omega_1,\omega_2,\omega_3\}. Then

{ω∈Ω∣D(ω)>100}={ω1,ω3}. \{\omega\in\Omega\mid D(\omega)>100\} =\{\omega_1,\omega_3\}.

This is the duration-threshold event in ℱ\mathcal F.

A value-space question becomes probabilistically meaningful by pointing back to an event in Ω\Omega.

Which mappings are allowed?

We need every value set we might ask about to point back to an observable event.

This requirement is measurability.

The value space also needs a rulebook

Let 𝒢\mathcal G be the collection of value sets we permit in probability questions.

Like the event space ℱ\mathcal F from Module 1, 𝒢\mathcal G must also permit:

  • the complement of every permitted set; and
  • the union of every finite or countably infinite list of permitted sets.

The structure ⟨𝒳,𝒢⟩\langle\mathcal X,\mathcal G\rangle is a measurable space.

Measurability makes value probabilities well defined

The function XX is measurable when

B∈𝒢⇒X−1(B)∈ℱ, B\in\mathcal G \quad\Longrightarrow\quad X^{-1}(B)\in\mathcal F,

Thus every permitted value question has an observable preimage event.

A random variable is a function X:⟨Ω,ℱ⟩→⟨𝒳,𝒢⟩X:\langle\Omega,\mathcal F\rangle\to\langle\mathcal X,\mathcal G\rangle satisfying this condition.

Notes: measurable random variables · Module 1: sigma-algebras

A case-coded pronoun variable

Let each outcome be a contextualized pronoun token. Define R(ω)R(\omega) as its combined form-and-case label, with codomain

(𝐼Nom,𝑚𝑒Acc,𝑦𝑜𝑢Nom,𝑦𝑜𝑢Acc,𝑡ℎ𝑒𝑦Nom,𝑡ℎ𝑒𝑚Acc,𝑖𝑡Nom,𝑖𝑡Acc,𝑠ℎ𝑒Nom,ℎ𝑒𝑟Acc,ℎ𝑒Nom,ℎ𝑖𝑚Acc,𝑤𝑒Nom,𝑢𝑠Acc). \begin{aligned} (&\textit{I}_{\mathrm{Nom}},\textit{me}_{\mathrm{Acc}}, \textit{you}_{\mathrm{Nom}},\textit{you}_{\mathrm{Acc}}, \textit{they}_{\mathrm{Nom}},\textit{them}_{\mathrm{Acc}},\\ &\textit{it}_{\mathrm{Nom}},\textit{it}_{\mathrm{Acc}}, \textit{she}_{\mathrm{Nom}},\textit{her}_{\mathrm{Acc}}, \textit{he}_{\mathrm{Nom}},\textit{him}_{\mathrm{Acc}}, \textit{we}_{\mathrm{Nom}},\textit{us}_{\mathrm{Acc}}). \end{aligned}

An integer may store a position in this ordered codomain, but that code has no linguistic meaning apart from the declared correspondence.

Preimages are events

For the first three labels in the declared order,

R−1({𝐼Nom,𝑚𝑒Acc,𝑦𝑜𝑢Nom})={ω∈Ω∣R(ω)∈{𝐼Nom,𝑚𝑒Acc,𝑦𝑜𝑢Nom}}. R^{-1}(\{\textit{I}_{\mathrm{Nom}},\textit{me}_{\mathrm{Acc}}, \textit{you}_{\mathrm{Nom}}\}) =\{\omega\in\Omega\mid R(\omega)\in \{\textit{I}_{\mathrm{Nom}},\textit{me}_{\mathrm{Acc}}, \textit{you}_{\mathrm{Nom}}\}\}.

More generally,

R−1(B)≡{ω∈Ω∣R(ω)∈B}. R^{-1}(B) \equiv\{\omega\in\Omega\mid R(\omega)\in B\}.

The measurability condition requires this preimage to belong to ℱ\mathcal F for every B∈𝒢B\in\mathcal G.

The set on the right contains token records, not three abstract pronoun spellings.

What moves from outcomes to values?

The probability space remains on Ω\Omega. The random variable induces probabilities for measurable sets of values.

Prediction: the probability of a value set BB must equal the probability of its preimage event.

Distribution of a random variable

The distribution of XX assigns each B∈𝒢B\in\mathcal G the probability

ℙ(X∈B)≡ℙ(X−1(B)). \mathbb{P}(X\in B) \equiv \mathbb{P}\!\left(X^{-1}(B)\right).

As a function of BB, this assignment is a probability measure on ⟨𝒳,𝒢⟩\langle\mathcal X,\mathcal G\rangle. We retain ℙ\mathbb{P} for the probability measure on ⟨Ω,ℱ⟩\langle\Omega,\mathcal F\rangle and introduce no second measure symbol.

Notes: reading probability through preimages

Can the possible values be listed?

This question identifies discrete distributions. We ask separately whether a non-discrete distribution has a density.

Discrete random variables

A random variable XX is discrete when some finite or countably infinite set 𝒮\mathcal S satisfies ℙ(X∈𝒮)=1\mathbb{P}(X\in\mathcal S)=1.

Annotated grammatical person has finite categorical support. A token count or an utterance-length count has finite or countably infinite numerical support.

Discrete does not mean numerical or unordered. It means that a listable set carries all probability.

Codomain and support

The codomain lists values the function may return. The support contains the values with positive probability:

supp⁡(X)≡{x∈cod⁡(X)∣ℙ(X=x)>0}. \operatorname{supp}(X) \equiv \{x\in\operatorname{cod}(X)\mid \mathbb P(X=x)>0\}.

Finite and countably infinite support

supp⁡(P)={first,second,third} \operatorname{supp}(P) =\{\mathrm{first},\mathrm{second},\mathrm{third}\}

is finite. An utterance-length support such as {0,1,2,…}\{0,1,2,\ldots\} is countably infinite. Both are discrete.

Values define events

L=5⇔{ω∈Ω∣L(ω)=5}. L=5 \quad\Longleftrightarrow\quad \{\omega\in\Omega\mid L(\omega)=5\}.

L≥5⇔{ω∈Ω∣L(ω)∈{5,6,7,…}}. L\geq5 \quad\Longleftrightarrow\quad \{\omega\in\Omega\mid L(\omega)\in\{5,6,7,\ldots\}\}.

An observed sample is not the support

For observed lengths

2,5,1,5,3, 2,5,1,5,3,

the distinct observed values are {1,2,3,5}\{1,2,3,5\}. The modeled support may still contain 00, 44, 66, and larger integers.

Discrete does not mean unordered

Grammatical person is discrete without a numerical distance between labels. Utterance length is discrete and ordered, with interpretable differences.

Discreteness concerns whether the probability-carrying values are separate and listable.

Notes: discrete support

How much probability belongs to each value?

For a discrete variable, separate codomain values can receive probability mass.

The function that records those probabilities is the probability mass function (PMF).

Defining a probability mass function

For a discrete random variable XX, its probability mass function (PMF) is defined by

pX:𝒳→[0,1],pX(x)≡ℙ(X=x). p_X:\mathcal X\to[0,1], \qquad p_X(x) \equiv\mathbb{P}(X=x).

The subscript names the random variable; the argument is a possible value.

Its support is the subset of codomain values with positive mass:

supp⁡(X)≡{x∈cod⁡(X):pX(x)>0}. \operatorname{supp}(X) \equiv \{x\in\operatorname{cod}(X):p_X(x)>0\}.

Notes: probability mass functions

Requirements for a PMF

Its PMF satisfies

pX(x)≥0and∑x∈cod⁡(X)pX(x)=1. p_X(x)\geq0 \qquad\text{and}\qquad \sum_{x\in\operatorname{cod}(X)}p_X(x)=1.

For any A∈𝒢A\in\mathcal G,

ℙ(X∈A)=∑x∈A∩cod⁡(X)pX(x). \mathbb{P}(X\in A) =\sum_{x\in A\cap\operatorname{cod}(X)}p_X(x).

Work a PMF question

Suppose SS counts syllables in a sampled word token.

syllable count ss 1 2 3 4
pS(s)p_S(s) .42 .33 .17 .08

Because {ω∈Ω∣S(ω)=3}\{\omega\in\Omega\mid S(\omega)=3\} and {ω∈Ω∣S(ω)=4}\{\omega\in\Omega\mid S(\omega)=4\} are disjoint,

ℙ(S≥3)=.17+.08=.25. \mathbb{P}(S\geq3)=.17+.08=.25.

For a discrete variable, an event probability is the sum of the masses at the codomain values in that event. Zero-mass values contribute zero.

Interpreting zero mass

If pS(7)=0p_S(7)=0, then

ℙ({ω∈Ω∣S(ω)=7})=0 \mathbb P\!\left(\{\omega\in\Omega\mid S(\omega)=7\}\right)=0

under the stated model. This does not imply that seven-syllable words are impossible in every linguistic population.

Omitted mass changes the model

Dropping the row with mass .08.08 leaves total mass .92.92, which is not a PMF.

Renormalizing the remaining rows defines a different model, conditional on at most three syllables.

What if modeled values fill an interval?

A duration is recorded at finite precision, but that storage convention does not force the model to have discrete support.

Prediction: interval probabilities should come from area, while individual points receive zero mass.

Which real-valued sets may receive probability?

A Borel set belongs to the smallest collection of subsets of ℝ\mathbb R that contains every open interval and is closed under complements and countable unions.

Thus ordinary intervals, threshold sets, and sets formed from them by these operations are Borel.

MIT probability notes: Borel sets

Absolutely continuous random variables

A real-valued random variable XX is absolutely continuous when there is a function fX:ℝ→[0,∞)f_X:\mathbb R\to[0,\infty) such that, for every Borel set A⊆ℝA\subseteq\mathbb R,

ℙ(X∈A)=∫AfX(x)dx. \mathbb{P}(X\in A)=\int_A f_X(x)\,\mathrm{d}x.

The function fXf_X is a probability density function.

Notes: continuous random variables

An uncountable codomain is not enough

Recall from Module 1 that a formant-frequency estimate records the estimated center frequency of a vocal-tract resonance. Suppose XX returns the first two such estimates ⟨F1,F2⟩\langle F_1,F_2\rangle from a vowel token. Let the idealized value space be ℝ+2\mathbb R_+^2 and define

X:Ω→ℝ+2,X(ω)≡⟨F1(ω),F2(ω)⟩. X:\Omega\to\mathbb R_+^2, \qquad X(\omega)\equiv\langle F_1(\omega),F_2(\omega)\rangle.

What absolute continuity adds

The codomain ℝ+2\mathbb R_+^2 is uncountable. That fact alone does not make the distribution absolutely continuous.

Absolute continuity requires probabilities to be recovered by integrating a density over ordinary two-dimensional area.

The units are hertz; the pair is an acoustic measurement, not a phoneme or vowel category.

A point and an interval are different events

Under a continuous model,

ℙ({ω∈Ω∣V(ω)=18.417})=0, \mathbb P\!\left(\{\omega\in\Omega\mid V(\omega)=18.417\}\right)=0,

while an interval around that point may receive positive probability:

ℙ({ω∈Ω∣18<V(ω)<19})>0. \mathbb P\!\left(\{\omega\in\Omega\mid 18<V(\omega)<19\}\right)>0.

Measurement precision does not choose the support

A stored value of 18 ms may be a rounded observation of an interval-valued quantity, or it may be the discrete output of the measurement process.

The scientific target determines which representation is appropriate.

Discrete and continuous support

question discrete continuous
values separate and listable fill a region
one value may have positive mass has probability zero
event probability add masses integrate density

Praat: formant measurements

What makes a function a PDF?

The density of XX satisfies

fX(x)≥0and∫−∞∞fX(x)dx=1. f_X(x)\geq0 \qquad\text{and}\qquad \int_{-\infty}^{\infty}f_X(x)\,\mathrm{d}x=1.

Conversely, any measurable nonnegative function with integral one defines an absolutely continuous probability distribution by integration.

Notes: probability density functions

Work an interval probability

Suppose

fD(d)≡{1/100,50≤d≤150,0,otherwise. f_D(d)\equiv \begin{cases} 1/100,&50\leq d\leq150,\\ 0,&\text{otherwise}. \end{cases}

Then

ℙ(80≤D≤100)=∫801001100du=.20. \mathbb{P}(80\leq D\leq100)=\int_{80}^{100}\frac{1}{100}\,\mathrm{d}u=.20.

This is a model for a measured token duration. It is not a claim that durations are actually uniform.

Density has reciprocal measurement units

If RR is measured in syllables per second, then fR(r)f_R(r) is measured in seconds per syllable.

Interval width times density is unitless probability.

A density height may exceed one

Height 2 across an interval of width .5.5 has total area

2(.5)=1. 2(.5)=1.

Density height is not bounded by one. Probability is.

Points have no mass under a density

For an absolutely continuous random variable,

ℙ(X=x)=∫xxfX(t)dt=0. \mathbb{P}(X=x)=\int_x^x f_X(t)\,\mathrm{d}t=0.

Thus fX(x)>0f_X(x)>0 does not imply ℙ(X=x)>0\mathbb{P}(X=x)>0.

Density height is not point probability. Only an integral of density gives probability.

Can one representation cover both cases?

Both a PMF and a PDF answer probability questions, but they do so differently.

The cumulative distribution function (CDF) instead asks one shared threshold question: how much probability lies at or below xx?

Defining a cumulative distribution function

Every real-valued random variable has a cumulative distribution function (CDF), defined by

FX:ℝ→[0,1],FX(x)≡ℙ(X≤x). F_X:\mathbb R\to[0,1], \qquad F_X(x) \equiv\mathbb{P}(X\leq x).

The same definition applies to discrete, continuous, and mixed distributions.

Notes: cumulative distribution functions

Compute the same object two ways

For a discrete XX,

FX(x)=∑t≤xpX(t). F_X(x)=\sum_{t\leq x}p_X(t).

For an absolutely continuous XX,

FX(x)=∫−∞xfX(t)dt. F_X(x)=\int_{-\infty}^{x}f_X(t)\,\mathrm{d}t.

These are consequences of the CDF definition.

PMFs, PDFs, and CDFs are different representations of distributions, not interchangeable names.

Mass, density, and accumulation

A PMF assigns mass to values. A PDF assigns density whose integral is probability. A CDF gives the probability at or below a threshold.

Expectations and variation

September 16

Which single number marks a distribution’s balance point?

A plain average gives equal weight to observed rows. A distributional average weights every possible value by its probability.

This probability-weighted average is the expected value.

When is the average well defined?

A variable is integrable when its expected absolute value is finite:

𝔼[|X|]<∞. \mathbb E[|X|]<\infty.

This requirement prevents undefined cancellation between infinite positive and negative contributions.

Expectations average with respect to probability

For an integrable random variable XX, define

𝔼[X]≡{∑x∈cod⁡(X)xpX(x),X discrete,∫−∞∞xfX(x)dx,X absolutely continuous. \mathbb E[X] \equiv \begin{cases} \sum_{x\in\operatorname{cod}(X)} x\,p_X(x),&X\text{ discrete},\\[2mm] \int_{-\infty}^{\infty}x f_X(x)\,\mathrm{d}x,&X\text{ absolutely continuous}. \end{cases}

Notes: expected values

Work a weighted average

Suppose SS counts syllables in a sampled word token and takes values 11, 22, and 44 with masses .25.25, .50.50, and .25.25. Then

𝔼[S]=1(.25)+2(.50)+4(.25)=2.25. \begin{aligned} \mathbb E[S] &=1(.25)+2(.50)+4(.25)\\ &=2.25. \end{aligned}

An expectation need not be a possible syllable count for any single token.

𝔼[S]=2.25\mathbb E[S]=2.25 describes the distribution, not an individual word token.

What if the target is not the original value?

Some questions concern a transformed value, such as squared distance from a reference point.

Prediction: transform each possible value first, then average those transformed values using the original probabilities.

Expectations of functions

When g(X)g(X) is integrable, define

𝔼[g(X)]≡∑x∈cod⁡(X)g(x)pX(x) \mathbb E[g(X)] \equiv\sum_{x\in\operatorname{cod}(X)}g(x)p_X(x)

in the discrete case, with the analogous integral in the continuous case.

In general, 𝔼[g(X)]≠g(𝔼[X])\mathbb E[g(X)]\neq g(\mathbb E[X]).

Notes: expectations of functions

Transform first, then average

For L∈{1,2,4}L\in\{1,2,4\} with masses (.25,.50,.25)(.25,.50,.25) and g(l)=l2g(l)=l^2,

𝔼[L2]=12(.25)+22(.50)+42(.25)=6.25. \mathbb E[L^2] =1^2(.25)+2^2(.50)+4^2(.25) =6.25.

Averaging first gives a different value

𝔼[L]=2.25but𝔼[L]2=5.0625. \mathbb E[L]=2.25 \qquad\text{but}\qquad \mathbb E[L]^2=5.0625.

Thus 𝔼[L2]≠𝔼[L]2\mathbb E[L^2]\neq\mathbb E[L]^2. The function encodes the property being averaged.

Which transformations can move through expectation?

For a linear transformation g(x)=ax+bg(x)=ax+b, averaging then transforming gives the same result as transforming then averaging.

Let us derive that claim before using it.

Indicators make the rule concrete

If IiI_i marks a target event and πi≡pIi(1)\pi_i\equiv p_{I_i}(1), then

𝔼[Ii]=πi. \mathbb E[I_i]=\pi_i.

For probabilities .20.20, .50.50, and .80.80,

𝔼[I1+I2+I3]=1.50. \mathbb E[I_1+I_2+I_3]=1.50.

Deriving linearity for a finite discrete model

For constants aa and bb,

𝔼[aX+b]=∑x∈cod⁡(X)(ax+b)pX(x)=a∑x∈cod⁡(X)xpX(x)+b∑x∈cod⁡(X)pX(x)=a𝔼[X]+b. \begin{aligned} \mathbb E[aX+b] &=\sum_{x\in\operatorname{cod}(X)}(ax+b)p_X(x)\\ &=a\sum_{x\in\operatorname{cod}(X)}xp_X(x) +b\sum_{x\in\operatorname{cod}(X)}p_X(x)\\ &=a\mathbb E[X]+b. \end{aligned}

Linearity for sums

If each XiX_i is integrable,

𝔼[∑i=1naiXi]=∑i=1nai𝔼[Xi]. \mathbb E\!\left[\sum_{i=1}^{n}a_iX_i\right] =\sum_{i=1}^{n}a_i\mathbb E[X_i].

This equality does not require independence. Dependence can change the distribution and variance of the sum.

Notes: linearity of expectation

Expectation adds across variables even when their values co-vary.

Independence is not required

The indicator values may be dependent. Linearity still gives

𝔼[I1+I2+I3]=𝔼[I1]+𝔼[I2]+𝔼[I3]. \mathbb E[I_1+I_2+I_3] =\mathbb E[I_1]+\mathbb E[I_2]+\mathbb E[I_3].

The corresponding product rule does require additional conditions:

𝔼[XY]≠𝔼[X]𝔼[Y] \mathbb E[XY]\neq\mathbb E[X]\mathbb E[Y]

in general.

Does every distribution have a mean?

No. The positive and negative parts of the weighted integral must not both diverge.

Symmetry does not guarantee an expectation

For X∼Cauchy⁡(x0,γ)X\sim\operatorname{Cauchy}(x_0,\gamma) with γ>0\gamma>0,

fX(x)≡1πγ[1+(x−x0γ)2]. f_X(x) \equiv \frac{1}{\pi\gamma\left[1+\left(\frac{x-x_0}{\gamma}\right)^2\right]}.

Why symmetric cancellation does not repair the mean

The positive and negative parts of ∫−∞∞xfX(x)dx\int_{-\infty}^{\infty}x f_X(x)\,\mathrm dx are both infinite. Thus 𝔼[X]\mathbb E[X] is undefined.

Integrating over [−M,M][-M,M] gives zero for every MM. The limit of these symmetric truncations is the symmetric principal value, not an expectation.

Notes: when expectations fail

Finite samples do not create a model expectation

Every finite sample of Cauchy draws has a finite arithmetic mean. An extreme draw can still move that mean by an arbitrarily large amount as the sample grows.

The model supplies no finite population mean to which the sample means must converge.

How far do values tend to lie from the mean?

Signed deviations cancel. Squaring them yields nonnegative distances and weights large deviations more heavily.

Their expectation is the variance.

Equal means can hide different spreads

Let D1D_1 equal 2 with probability one. Let D2D_2 equal 1 or 3, each with probability .50.50.

𝔼[D1]=𝔼[D2]=2,Var⁡(D1)=0,Var⁡(D2)=1. \mathbb E[D_1]=\mathbb E[D_2]=2, \qquad \operatorname{Var}(D_1)=0, \qquad \operatorname{Var}(D_2)=1.

What must exist before variance?

The second moment is 𝔼[X2]\mathbb E[X^2].

It is finite when

𝔼[X2]<∞. \mathbb E[X^2]<\infty.

This condition ensures that the squared deviations used by variance have a finite average.

Defining variance

For a random variable with a finite second moment, define

Var⁡(X)≡𝔼[(X−𝔼[X])2]. \operatorname{Var}(X) \equiv\mathbb E\!\left[(X-\mathbb E[X])^2\right].

Variance is an expected squared deviation from the mean.

Notes: variance

Derive an equivalent calculation

Writing μ≡𝔼[X]\mu\equiv\mathbb E[X],

Var⁡(X)=𝔼[X2−2μX+μ2]=𝔼[X2]−2μ𝔼[X]+μ2=𝔼[X2]−2μ2+μ2=𝔼[X2]−μ2=𝔼[X2]−𝔼[X]2. \begin{aligned} \operatorname{Var}(X) &=\mathbb E[X^2-2\mu X+\mu^2]\\ &=\mathbb E[X^2]-2\mu\mathbb E[X]+\mu^2\\ &=\mathbb E[X^2]-2\mu^2+\mu^2\\ &=\mathbb E[X^2]-\mu^2\\ &=\mathbb E[X^2]-\mathbb E[X]^2. \end{aligned}

Standard deviation

Variance squares the measurement unit. Taking its square root returns to the original unit.

The standard deviation is defined by

σX≡Var⁡(X). \sigma_X \equiv\sqrt{\operatorname{Var}(X)}.

If XX is measured in milliseconds, Var⁡(X)\operatorname{Var}(X) has units ms2^2 and σX\sigma_X has units ms.

Notes: standard deviation

Voice onset time on the original scale

If

Var⁡(V)=625 ms2, \operatorname{Var}(V)=625\text{ ms}^2,

then

σV=625 ms2=25 ms. \sigma_V=\sqrt{625\text{ ms}^2}=25\text{ ms}.

Standard deviation is not a complete shape description

Standard deviation is the root mean squared distance from the mean, not the average absolute distance.

Equal means and standard deviations do not guarantee equal tails, symmetry, or modality.

Standard deviation does not determine coverage

The mean and standard deviation do not determine how much probability lies within μ±σ\mu\pm\sigma.

Coverage claims require an additional distribution-shape assumption.

What else can centered powers retain?

Variance uses the second power. Higher centered powers can summarize other aspects of distribution shape, when the corresponding expectations exist.

Defining central moments

When it exists, the kkth central moment is defined by

μk≡𝔼[(X−𝔼[X])k]. \mu_k \equiv\mathbb E\!\left[(X-\mathbb E[X])^k\right].

Thus μ1=0\mu_1=0 and μ2=Var⁡(X)\mu_2=\operatorname{Var}(X).

Notes: central moments

Higher central moments

For J∈{1,2,5}J\in\{1,2,5\} with masses (.25,.50,.25)(.25,.50,.25) and 𝔼[J]=2.5\mathbb E[J]=2.5,

μ3=3,μ4=11.0625. \mu_3=3, \qquad \mu_4=11.0625.

Odd powers retain direction. Higher powers give more weight to values far from the mean.

Moments do not identify a distribution

Different distributions can share the same first several moments.

Thus moments summarize specified centered powers; they do not determine the observation process or the full distributional shape.

An expectation inherits its function and units

An expectation summarizes a specified function of a distribution. Its meaning depends on that function, the existence of the expectation, and the measurement units.

Standard distribution families

September 21

What does choosing a family assert?

A distribution family links a support to a particular probability form.

Thus family choice must follow the observation process and the research question, not merely the shape of a histogram.

Distribution notation

A parameter is a fixed value that selects one distribution from a family. Changing it may change location, spread, shape, or several of these properties.

The statement

X∼𝒟(θ) X\sim\mathcal D(\theta)

means that the distribution of XX belongs to family 𝒟\mathcal D with parameter θ\theta. The symbol ∼\sim does not mean numerical equality.

It says that XX has a distribution in the named family under the stated parameter convention.

Question: which category does one token receive?

Suppose one contextualized pronoun token receives exactly one form-and-case label from a finite list.

The categorical distribution assigns one probability to each label.

The categorical distribution

For labels c1,…,cKc_1,\ldots,c_K and 𝝅∈[0,1]K\boldsymbol\pi\in[0,1]^K with ∑k=1Kπk=1\sum_{k=1}^K\pi_k=1,

R∼Categorical⁡(𝝅) R\sim\operatorname{Categorical}(\boldsymbol\pi)

means

pR(ck)≡ℙ(R=ck)=πk. p_R(c_k) \equiv\mathbb{P}(R=c_k) =\pi_k.

Notes: categorical distributions · Stan reference

Declare the category order

Order the pronoun labels as shown in the graph and define the probability vector

𝝅≡(.03,.09,.03,.12,.07,.28,.07,.05,.02,.02,.07,.08,.05,.02). \boldsymbol\pi \equiv(.03,.09,.03,.12,.07,.28,.07,.05,.02,.02,.07,.08,.05,.02).

The kkth entry is the probability of the kkth form-and-case label. Reordering the labels without applying the same permutation to 𝝅\boldsymbol\pi changes the distribution.

The labels encode two annotated attributes. They are not fourteen distinct English spellings.

PMF for the pronoun categorical distribution

Question: does the token have Case=Acc?

The categorical record retains all fourteen labels. If the question concerns one binary contrast, we can map those labels to 1 and 0.

Prediction: different spellings may receive the same value, and a syncretic spelling may receive different values in different contexts.

Define the case indicator

Define the accusative event

AAcc≡{ω∈Ω∣C(ω)=Acc}. A_{\mathrm{Acc}}\equiv\{\omega\in\Omega\mid C(\omega)=\mathrm{Acc}\}.

Define the accusative-case indicator

X(ω)≡{1,ω∈AAcc,0,ω∉AAcc. X(\omega) \equiv \begin{cases} 1,&\omega\in A_{\mathrm{Acc}},\\ 0,&\omega\notin A_{\mathrm{Acc}}. \end{cases}

The declared codomain is {0,1}\{0,1\} even though the sample space contains fourteen pronoun outcomes. For 0≤π≤10\leq\pi\leq1,

X∼Bernoulli⁡(π) X\sim\operatorname{Bernoulli}(\pi)

means

pX(x)≡ℙ(X=x)=πx(1−π)1−x,x∈{0,1}. p_X(x) \equiv\mathbb{P}(X=x) =\pi^x(1-\pi)^{1-x}, \qquad x\in\{0,1\}.

Notes: Bernoulli distributions · UD English Case

Bernoulli distribution on pronoun case

Using π=.27\pi=.27,

pX(1)=.27,pX(0)=.73. p_X(1)=.27, \qquad p_X(0)=.73.

The parameter is the probability of annotated accusative case because the definition of XX assigns Case=Acc outcomes the value 11.

Bernoulli parameters inherit their meaning from the event coded as 1.

Bernoulli mean and variance

For X∼Bernoulli⁡(π)X\sim\operatorname{Bernoulli}(\pi),

𝔼[X]=π,Var⁡(X)=π(1−π). \mathbb E[X]=\pi, \qquad \operatorname{Var}(X)=\pi(1-\pi).

The mean equals the probability of the event coded as 1.

The coding determines the parameter

If the coding is reversed, the parameter becomes the probability of the complementary event.

Thus π\pi is not an intrinsic property of the spelling. It inherits its meaning from the declared outcome-to-value mapping.

Question: how many successes occur in fixed opportunities?

Suppose we predeclare nn token opportunities and count how many have Case=Acc.

Prediction: a binomial model needs a fixed nn, independent indicators, and one common success probability π\pi.

The binomial distribution

Let Y1,…,YnY_1,\ldots,Y_n be independent Bernoulli variables with common parameter π\pi, and define K≡∑i=1nYiK\equiv\sum_{i=1}^nY_i. Then

K∼Binomial⁡(n,π), K\sim\operatorname{Binomial}(n,\pi),

with support {0,…,n}\{0,\ldots,n\}.

Notes: binomial distributions

Why does the binomial coefficient appear?

The binomial coefficient (nk){n\choose k} counts the ways to choose kk positions from nn. Each sequence with kk ones has probability πk(1−π)n−k\pi^k(1-\pi)^{n-k}. Thus

pK(k)≡ℙ(K=k)=(nk)πk(1−π)n−k. p_K(k) \equiv\mathbb{P}(K=k) ={n\choose k}\pi^k(1-\pi)^{n-k}.

(nk){n\choose k} counts the distinct indicator sequences that produce the same total kk.

A binomial probability by hand

For n=6n=6, π=.25\pi=.25, and k=2k=2,

pK(2)=(62)(.25)2(.75)4≈.297. p_K(2) ={6\choose2}(.25)^2(.75)^4 \approx.297.

The calculation concerns exactly two target outcomes among six declared opportunities.

Binomial center and spread

For K∼Binomial⁡(n,π)K\sim\operatorname{Binomial}(n,\pi),

𝔼[K]=nπ,Var⁡(K)=nπ(1−π). \mathbb E[K]=n\pi, \qquad \operatorname{Var}(K)=n\pi(1-\pi).

The linguistic assumptions remain visible

The model requires a fixed number of opportunities, independent indicators, and one common success probability.

The count K=2K=2 does not retain which two positions were successes, and the same proportion can arise from different values of nn.

What changes without replacement?

If selecting one clause changes which clauses remain available, later draws are not independent Bernoulli trials.

The hypergeometric distribution represents a count from a fixed finite population sampled without replacement.

Sampling without replacement

Suppose an archive contains NN clauses, MM of which are passive. Select nn clauses without replacement. If KK counts passives, then

pK(k)≡ℙ(K=k)=(Mk)(N−Mn−k)(Nn). p_K(k) \equiv\mathbb{P}(K=k) =\frac{{M\choose k}{N-M\choose n-k}}{{N\choose n}}.

The support satisfies

max⁡(0,n−N+M)≤k≤min⁡(n,M). \max(0,n-N+M)\leq k\leq\min(n,M).

The linguistic unit

One unit is a clause record, and passive voice is a declared annotation on that record.

Under English Universal Dependencies, Voice=Pass marks a verbal past participle used as the predicate of a passive clause.

Notes: hypergeometric distributions · UD English Voice · R manual

A hypergeometric probability by hand

For N=30N=30, M=9M=9, n=5n=5, and k=2k=2,

pK(2)=(92)(213)(305)≈.336. p_K(2) =\frac{{9\choose2}{21\choose3}}{{30\choose5}} \approx.336.

Mean and finite population correction

𝔼[K]=5930=1.5,Var⁡(K)=5(.3)(.7)2529. \mathbb E[K]=5\frac9{30}=1.5, \qquad \operatorname{Var}(K) =5(.3)(.7)\frac{25}{29}.

The factor 25/2925/29 reduces variance because sampling one clause changes what remains.

Without replacement versus independent draws

A binomial count from independent draws with probability .30.30 also has mean 1.51.5, but variance 5(.3)(.7)5(.3)(.7).

The hypergeometric variance is smaller because the draws are dependent after clauses are removed.

Question: how long do we wait for one target event?

Let each opportunity be a Bernoulli trial, and let KK count failures before the first success.

Before naming the family, we need probabilities over 0,1,2,…0,1,2,\ldots that sum to one.

A geometric series defines a PMF

A geometric series multiplies each successive term by the same ratio. Begin with

∑k=0∞12k+1=1. \sum_{k=0}^{\infty}\frac{1}{2^{k+1}}=1.

Thus

pK(k)≡12k+1,k=0,1,2,… p_K(k)\equiv\frac{1}{2^{k+1}}, \qquad k=0,1,2,\ldots

is a PMF with countably infinite support.

Define the geometric family

Let KK count failures before the first success in independent Bernoulli trials with common success probability π\pi. Define

K∼Geometric⁡(π) K\sim\operatorname{Geometric}(\pi)

by the PMF

pK(k)≡ℙ(K=k)=(1−π)kπ,k=0,1,2,… p_K(k) \equiv\mathbb{P}(K=k) =(1-\pi)^k\pi, \qquad k=0,1,2,\ldots

This is the convention used by dgeom in R. The total-trial count is T≡K+1T\equiv K+1.

Notes: geometric distributions · R manual: geometric distribution

A geometric waiting probability

For π=.20\pi=.20,

pK(3)=(.80)3(.20)=.1024,ℙ(K≥4)=(.80)4=.4096. p_K(3)=(.80)^3(.20)=.1024, \qquad \mathbb P(K\geq4)=(.80)^4=.4096.

The first quantity fixes the first target after three failures; the second requires four initial failures.

Deriving the geometric expectation

Let q≡1−πq\equiv1-\pi. Differentiating the geometric series gives

∑k=0∞qk=11−q⇒∑k=1∞kqk−1=1(1−q)2. \sum_{k=0}^{\infty}q^k=\frac{1}{1-q} \quad\Longrightarrow\quad \sum_{k=1}^{\infty}kq^{k-1}=\frac{1}{(1-q)^2}.

Thus

𝔼[K]=∑k=0∞kqkπ=πq∑k=1∞kqk−1=1−ππ. \begin{aligned} \mathbb E[K] &=\sum_{k=0}^{\infty}kq^k\pi\\ &=\pi q\sum_{k=1}^{\infty}kq^{k-1}\\ &=\frac{1-\pi}{\pi}. \end{aligned}

Geometric⁡(.5)\operatorname{Geometric}(.5)

As π\pi decreases, the distribution spreads rightward

As π\pi increases, mass concentrates at zero

Every geometric PMF decreases from its first support point

The mode is a value with the greatest probability mass. A mode at the smallest possible value is a boundary mode.

For every k≥0k\geq0,

pK(k+1)pK(k)=(1−π)k+1π(1−π)kπ=1−π<1. \frac{p_K(k+1)}{p_K(k)} =\frac{(1-\pi)^{k+1}\pi}{(1-\pi)^k\pi} =1-\pi <1.

Thus pK(k+1)<pK(k)p_K(k+1)<p_K(k). If a positive count is L≡K+1L\equiv K+1, every geometric model has its mode at L=1L=1.

This is a model prediction. We can now compare it with an empirical distribution.

What family removes the boundary-mode restriction?

The negative binomial distribution counts failures before the rrth success. The additional parameter rr permits an interior mode.

This is a family comparison, not yet a claim that dictionary symbols arise from literal Bernoulli trials.

The negative binomial generalizes the geometric

Let KK count failures before the rrth success, where r∈{1,2,…}r\in\{1,2,\ldots\}. Define

K∼NegBin⁡(r,π) K\sim\operatorname{NegBin}(r,\pi)

by

pK(k)≡ℙ(K=k)=(k+r−1k)(1−π)kπr,k=0,1,2,… p_K(k) \equiv\mathbb{P}(K=k) ={k+r-1\choose k}(1-\pi)^k\pi^r, \qquad k=0,1,2,\ldots

When r=1r=1, the binomial coefficient equals one and this PMF is exactly geometric.

Notes: negative-binomial distributions · R manual

Begin with r=1r=1 and π=.5\pi=.5

Increasing rr permits an interior mode

With π\pi fixed, increasing rr increases

𝔼[K]=r1−ππ. \mathbb E[K]=r\frac{1-\pi}{\pi}.

With r=5r=5, smaller π\pi moves mass farther right

With r=5r=5, larger π\pi concentrates mass near zero

A larger rr can offset a large π\pi

Mean and variance under the negative binomial

Set K≡L−1K\equiv L-1 so that the observed support begins at zero. Under this negative-binomial convention,

𝔼[K]=r1−ππ,Var⁡(K)=r1−ππ2. \mathbb E[K]=r\frac{1-\pi}{\pi}, \qquad \operatorname{Var}(K)=r\frac{1-\pi}{\pi^2}.

What does dispersion mean for a count?

For a count distribution, dispersion compares variance with the mean.

  • Overdispersion: Var⁡(K)>𝔼[K]\operatorname{Var}(K)>\mathbb E[K].
  • Equidispersion: Var⁡(K)=𝔼[K]\operatorname{Var}(K)=\mathbb E[K].
  • Underdispersion: Var⁡(K)<𝔼[K]\operatorname{Var}(K)<\mathbb E[K].

What dispersion does the family predict?

Writing μ≡𝔼[K]\mu\equiv\mathbb E[K] gives

Var⁡(K)=μ+μ2r>μ \operatorname{Var}(K)=\mu+\frac{\mu^2}{r}>\mu

for every finite rr.

Thus every finite negative-binomial model is overdispersed relative to its mean.

A negative-binomial probability by hand

For r=3r=3, π=.20\pi=.20, and k=4k=4,

pK(4)=(64)(.80)4(.20)3=.049152. p_K(4) ={6\choose4}(.80)^4(.20)^3 =.049152.

The final draw is the third target; the four failures occupy the first six positions.

Model probabilities or observed proportions?

A model PMF assigns probabilities to possible values.

An empirical distribution assigns each observed value its relative frequency.

CMUdict pronunciation-entry length

Use every noncomment pronunciation entry in CMU Pronouncing Dictionary 0.7b. For entry ω\omega, define

L(ω)≡number of ARPAbet-based pronunciation symbols in ω. L(\omega)\equiv\text{number of ARPAbet-based pronunciation symbols in }\omega.

The unit is a pronunciation entry, not a corpus token or necessarily a unique spelling type.

CMUdict file-format documentation · Notes: the worked example

Pronunciation-entry lengths have an interior mode

The geometric family misses the observed mode

The empirical relative-frequency function has its mode at six symbols. A fitted geometric PMF must have its mode at one symbol.

How does likelihood compare parameter values?

Hold the observed data fixed. The likelihood is the joint PMF evaluated at observed discrete data, or the joint density evaluated at observed continuous data, treated as a function of the parameters.

A maximum-likelihood estimate is a parameter value that maximizes this function, when such a value exists. A larger likelihood means a better fit to these observations within the compared family.

Compare the prediction with CMUdict

Prediction: a finite rr cannot match underdispersed data, whose variance is below their mean.

For the pronunciation-entry lengths, the mean of KK is 5.381 and its variance is 4.566.

For each rr, estimate π\pi and evaluate the likelihood of all observed entry lengths.

Negative-binomial fit

The likelihood increases toward r→∞r\to\infty, the Poisson boundary derived next.

The extra parameter fixes the geometric mode restriction, but every finite negative-binomial model remains overdispersed relative to its mean.

Empirical, geometric, and Poisson-limit CDFs

The negative-binomial generalization removes the geometric family’s boundary-mode restriction. For these dictionary entries, the maximum-likelihood fit approaches its Poisson limit rather than a finite rr.

An added parameter can repair one mismatch while leaving another. Model criticism must name both.

Which family appears at the dispersion boundary?

The fitted negative-binomial sequence approaches a Poisson distribution while keeping its mean fixed.

Let us state that limit before interpreting a Poisson count.

The Poisson limiting case

Fix λ>0\lambda>0 and let

Kr∼NegBin⁡(r,rr+λ). K_r\sim\operatorname{NegBin}\!\left(r,\frac{r}{r+\lambda}\right).

Then, for every k=0,1,2,…k=0,1,2,\ldots,

limr→∞ℙ(Kr=k)=e−λλkk!. \lim_{r\to\infty}\mathbb P(K_r=k) =e^{-\lambda}\frac{\lambda^k}{k!}.

Thus the negative-binomial family approaches Poisson⁡(λ)\operatorname{Poisson}(\lambda) as r→∞r\to\infty while its mean remains λ\lambda.

Question: a count per what opportunity?

A count has no interpretable rate without an exposure: tokens, clauses, speakers, minutes, or another declared unit.

Poisson counts require a declared exposure

If events occur at rate λ>0\lambda>0 per exposure unit, a Poisson model for exposure t>0t>0 specifies

Yt∼Poisson⁡(λt), Y_t\sim\operatorname{Poisson}(\lambda t),

with

pYt(y)≡ℙ(Yt=y)=e−λt(λt)yy!,y=0,1,2,… p_{Y_t}(y) \equiv\mathbb{P}(Y_t=y) =e^{-\lambda t}\frac{(\lambda t)^y}{y!}, \qquad y=0,1,2,\ldots

Notes: Poisson distributions · R manual

Poisson probabilities by hand

For λt=2\lambda t=2,

pYt(0)=e−2≈.135,pYt(2)=e−2222!≈.271. p_{Y_t}(0)=e^{-2}\approx.135, \qquad p_{Y_t}(2)=\frac{e^{-2}2^2}{2!}\approx.271.

The first probability concerns no filled pauses; the second concerns exactly two in the declared exposure.

A Poisson distribution is not a complete process claim

The count model implies

𝔼[Yt]=Var⁡(Yt)=λt. \mathbb E[Y_t]=\operatorname{Var}(Y_t)=\lambda t.

A Poisson process with a constant rate additionally specifies N(0)=0N(0)=0, independent counts in disjoint intervals, and the same expected count for intervals of the same length.

A Poisson marginal distribution does not by itself license a story about a Poisson process.

Question: what if equal-length intervals receive equal probability?

On a finite interval, constant density produces probability proportional to interval length.

This is the continuous uniform distribution.

The continuous uniform distribution

For a<ba<b,

X∼Uniform⁡(a,b) X\sim\operatorname{Uniform}(a,b)

means

fX(x)≡{1/(b−a),a≤x≤b,0,otherwise. f_X(x)\equiv \begin{cases} 1/(b-a),&a\leq x\leq b,\\ 0,&\text{otherwise}. \end{cases}

Notes: uniform distributions · R manual

Density and probability under Uniform⁡(0,1)\operatorname{Uniform}(0,1)

For X∼Uniform⁡(0,1)X\sim\operatorname{Uniform}(0,1),

ℙ(.25<X<.75)=∫.25.751dx=.50. \mathbb P(.25<X<.75) =\int_{.25}^{.75}1\,\mathrm dx =.50.

Probability is area under the density

Equal density height means equal probability for intervals of equal length, not equal positive probability at every point.

CDF of Uniform⁡(0,1)\operatorname{Uniform}(0,1)

Linguistic use: randomization by design

Let U∼Uniform⁡(0,1)U\sim\operatorname{Uniform}(0,1) be a computer-generated draw used to assign a trial to one of two lists:

A(U)={list 1,U<.5,list 2,U≥.5. A(U)= \begin{cases} \text{list 1},&U<.5,\\ \text{list 2},&U\geq.5. \end{cases}

Here uniformity describes the randomization mechanism. It does not assert that an acoustic measurement is naturally uniform.

Question: how can a distribution vary within (0,1)(0,1)?

Proportions and bounded continuous responses may cluster near the center, near one boundary, or near both boundaries.

The beta family can represent these contrasting shapes through two positive parameters.

The beta distribution

For α,β>0\alpha,\beta>0, the beta function supplies the normalizing area

B(α,β)≡∫01tα−1(1−t)β−1dt. B(\alpha,\beta) \equiv\int_0^1 t^{\alpha-1}(1-t)^{\beta-1}\,\mathrm dt.

Then

X∼Beta⁡(α,β) X\sim\operatorname{Beta}(\alpha,\beta)

has possible values in [0,1][0,1] and density

fX(x)≡{xα−1(1−x)β−1B(α,β),0<x<1,0,otherwise. f_X(x)\equiv \begin{cases} \dfrac{x^{\alpha-1}(1-x)^{\beta-1}}{B(\alpha,\beta)},&0<x<1,\\ 0,&\text{otherwise}. \end{cases}

Notes: beta distributions · R manual

Beta⁡(1,1)\operatorname{Beta}(1,1) is uniform on (0,1)(0,1)

Symmetric beta densities with α=β>1\alpha=\beta>1

Asymmetric beta densities

Boundary modes in the beta family

Two boundary modes when α,β<1\alpha,\beta<1

Transforming beta support to (a,b)(a,b)

If X∼Beta⁡(α,β)X\sim\operatorname{Beta}(\alpha,\beta) and

Y≡a+(b−a)X, Y\equiv a+(b-a)X,

then, for a<y<ba<y<b,

fY(y)≡1b−a(y−ab−a)α−1(1−y−ab−a)β−1B(α,β). f_Y(y) \equiv \frac{1}{b-a} \frac{ \left(\frac{y-a}{b-a}\right)^{\alpha-1} \left(1-\frac{y-a}{b-a}\right)^{\beta-1} }{B(\alpha,\beta)}.

The change-of-scale factor 1/(b−a)1/(b-a) is required for the transformed density to integrate to one.

Interpreting beta parameters

For X∼Beta⁡(α,β)X\sim\operatorname{Beta}(\alpha,\beta),

𝔼[X]=1B(α,β)∫01xα(1−x)β−1dx=B(α+1,β)B(α,β)=αα+β. \begin{aligned} \mathbb E[X] &=\frac{1}{B(\alpha,\beta)} \int_0^1x^\alpha(1-x)^{\beta-1}\,\mathrm dx\\ &=\frac{B(\alpha+1,\beta)}{B(\alpha,\beta)}\\ &=\frac{\alpha}{\alpha+\beta}. \end{aligned}

Holding this ratio fixed while increasing α+β\alpha+\beta concentrates the distribution around the same mean.

α/(α+β)\alpha/(\alpha+\beta) controls the mean, while α+β\alpha+\beta controls concentration when the mean is held fixed.

Linguistic use: a bounded continuous response

Suppose AA is a participant’s sentence-acceptability rating, recorded with a continuous slider and rescaled to (0,1)(0,1).

A beta model treats AA as a measurement, not as a grammatical category. If the task permits exact endpoint responses 0 and 1, an ordinary beta distribution is not sufficient without an endpoint-handling component.

Question: how can location and scale vary independently?

The normal family uses one parameter for its center and another for its spread over the real line.

This support assumption matters: an untransformed duration or proportion cannot literally take every real value.

The normal distribution

For σ>0\sigma>0,

X∼𝒩(μ,σ2) X\sim\mathcal N(\mu,\sigma^2)

means

fX(x)≡1σ2πexp⁡[−(x−μ)22σ2]. f_X(x)\equiv\frac{1}{\sigma\sqrt{2\pi}} \exp\!\left[-\frac{(x-\mu)^2}{2\sigma^2}\right].

Notes: normal distributions · R manual

PDF of the standard normal distribution

The mean controls location; the standard deviation controls scale

Standardizing a normal variable

Standardization expresses a value’s distance from the model mean in standard-deviation units.

Define

Z≡X−μσ. Z\equiv\frac{X-\mu}{\sigma}.

If X∼𝒩(μ,σ2)X\sim\mathcal N(\mu,\sigma^2), then Z∼𝒩(0,1)Z\sim\mathcal N(0,1). Standardization removes the physical unit and retains relative position within this family.

A standardized value reports position in standard-deviation units under the assumed normal model.

A normal interval by hand

For X∼𝒩(550,402)X\sim\mathcal N(550,40^2),

510−55040=−1,590−55040=1. \frac{510-550}{40}=-1, \qquad \frac{590-550}{40}=1.

Thus ℙ(510≤X≤590)≈.683\mathbb P(510\leq X\leq590)\approx.683 under this normal model.

Linguistic use: log duration

Let D>0D>0 be a token duration and Y≡log⁡DY\equiv\log D. A normal model for YY states

Y∼𝒩(μ,σ2). Y\sim\mathcal N(\mu,\sigma^2).

This is a distributional assumption about a transformed acoustic measurement. It is not implied by the fact that a duration histogram looks roughly bell-shaped.

Standard normal CDF

A beta CDF can also be S-shaped

Question: what is the distribution of squared normal discrepancies?

Squaring standard-normal variables makes every contribution nonnegative. Adding ν\nu independent squared values yields a chi-squared variable.

The count ν\nu is called its degrees of freedom: it records the number of independent squared contributions.

The chi-squared distribution

For ν∈{1,2,…}\nu\in\{1,2,\ldots\}, let Z1,…,ZνZ_1,\ldots,Z_\nu be independent standard normal variables and define

Q≡∑i=1νZi2. Q\equiv\sum_{i=1}^{\nu}Z_i^2.

Then Q∼χν2Q\sim\chi^2_\nu, with 𝔼[Q]=ν\mathbb E[Q]=\nu and Var⁡(Q)=2ν\operatorname{Var}(Q)=2\nu.

Notes: chi-squared distributions

Linguistic use: an aggregate discrepancy

Suppose Z1,Z2,Z3Z_1,Z_2,Z_3 are independent standardized discrepancies for three predeclared acoustic measurements. Then

Q=Z12+Z22+Z32∼χ32 Q=Z_1^2+Z_2^2+Z_3^2\sim\chi^2_3

under the standard-normal assumptions. The degrees of freedom count independent squared contributions, not speech tokens.

A chi-squared discrepancy by hand

For Z1=1Z_1=1, Z2=−.5Z_2=-.5, and Z3=2Z_3=2,

Q=12+(−.5)2+22=5.25. Q=1^2+(-.5)^2+2^2=5.25.

Under the independent standard-normal construction, Q∼χ32Q\sim\chi^2_3 and ℙ(Q≥5.25)≈.154\mathbb P(Q\geq5.25)\approx.154.

Chi-squared center and spread

For Q∼χν2Q\sim\chi^2_\nu,

𝔼[Q]=ν,Var⁡(Q)=2ν. \mathbb E[Q]=\nu, \qquad \operatorname{Var}(Q)=2\nu.

Degrees of freedom determine both center and spread.

An upper-tail probability

For the constructed value Q=5.25Q=5.25 with ν=3\nu=3,

ℙ(Q≥5.25)≈.154. \mathbb P(Q\geq5.25)\approx.154.

This is probability under the reference distribution, not the probability that a hypothesis is true.

Degrees of freedom count unconstrained variation

In an rr by cc table with fixed row and column totals,

ν=(r−1)(c−1). \nu=(r-1)(c-1).

The symbol ν\nu counts independent pieces of variation remaining after constraints, not necessarily observations.

What changes when scale is estimated?

Dividing a standard-normal variable by an independent estimated scale adds uncertainty. The resulting Student’s t distribution has heavier tails, that is, it places more probability far from zero.

Student’s t distribution

Let Z∼𝒩(0,1)Z\sim\mathcal N(0,1) and U∼χν2U\sim\chi^2_\nu be independent. Define

T≡ZU/ν. T\equiv\frac{Z}{\sqrt{U/\nu}}.

Then T∼tνT\sim t_\nu. Because U/νU/\nu can be close to zero, extreme values of TT are more probable than extreme values under a standard normal distribution.

Notes: Student’s t distributions · R manual

A two-sided t probability

For |t|=2.5|t|=2.5 and ν=5\nu=5,

ℙ(|T|≥2.5)=2ℙ(T≥2.5)≈.0545. \mathbb P(|T|\geq2.5) =2\mathbb P(T\geq2.5) \approx.0545.

The corresponding standard-normal probability is about .0124.0124 because t5t_5 has heavier tails.

Linguistic use: keep two roles separate

An estimator is a rule that maps a sample to an estimate of an unknown parameter. A tt distribution may describe a standardized estimator whose scale was estimated.

It may instead be chosen as an observation model, that is, a distribution for a transformed linguistic response.

Those are different claims. Naming the random variable and its generating construction keeps them apart.

A distribution family is not a process model

A distribution family specifies a support and a mass or density function. A process interpretation requires additional assumptions about how observations are generated.

How can a CDF generate draws?

A CDF partitions probability from zero to one. A uniform draw can select one of those probability intervals.

This is the intuition behind inverse transform sampling.

Inverse transform sampling

For inverse transform sampling, the generalized inverse returns the leftmost value at which the CDF reaches uu. Formally,

F−1(u)≡inf⁡{x∈ℝ:F(x)≥u},0<u<1. F^{-1}(u) \equiv\inf\{x\in\mathbb R:F(x)\geq u\}, \qquad 0<u<1.

If U∼Uniform⁡(0,1)U\sim\operatorname{Uniform}(0,1), define X≡F−1(U)X\equiv F^{-1}(U).

Notes: inverse transform sampling

Verifying the target distribution

The generalized inverse satisfies F−1(U)≤xF^{-1}(U)\leq x exactly when U≤F(x)U\leq F(x). Thus

ℙ(X≤x)=ℙ(F−1(U)≤x)=ℙ(U≤F(x))=F(x). \begin{aligned} \mathbb{P}(X\leq x) &=\mathbb{P}(F^{-1}(U)\leq x)\\ &=\mathbb{P}(U\leq F(x))\\ &=F(x). \end{aligned}

Discrete inverse transform sampling

For ordered values x1,…,xKx_1,\ldots,x_K, define

ck≡∑j=1kpX(xj). c_k\equiv\sum_{j=1}^kp_X(x_j).

Draw U∼Uniform⁡(0,1)U\sim\operatorname{Uniform}(0,1) and return the first xkx_k for which U≤ckU\leq c_k.

Cumulative mass turns one uniform value into one draw from the target PMF.

Cumulative cutpoints

label mass endpoint interval
Nom .50.50 .50.50 (0,.50](0,.50]
Acc .30.30 .80.80 (.50,.80](.50,.80]
Gen .20.20 11 (.80,1)(.80,1)

Work a categorical draw

Suppose a constructed three-label PMF assigns (.50,.30,.20)(.50,.30,.20) to (Nom, Acc, Gen). Its cumulative cutpoints are (.50,.80,1)(.50,.80,1).

If U=.73U=.73, return Acc, since .50<U≤.80.50<U\leq.80.

The returned value is an annotated category label, not a spelling or pronunciation.

Checking the implemented cutpoints

The endpoints must be the cumulative sums (.50,.80,1)(.50,.80,1).

If the first endpoint were .40.40, the first label would be simulated too rarely. Large-sample proportions would expose the error.

What does a random seed guarantee?

A pseudorandom algorithm is deterministic once its initial state is fixed.

A random seed records that initial state.

Random seeds reproduce computations

A random seed fixes the initial state of a pseudorandom number generator. Reusing the seed reproduces the same sequence of draws.

It does not validate the distribution, independence assumptions, or sampling representation.

Notes: random seeds · R manual: Random

One seed, one computational sequence

Set the seed once before a related block of random draws. Resetting the same seed inside a loop repeats the same simulated result.

Record the seed with the code and place it before the first draw that affects the reported result.

Stability across seeds

Changing the seed provides a check on Monte Carlo variation. A conclusion that changes materially across seeds may require more draws or an implementation check.

More draws reduce simulation variation. They do not repair the wrong target distribution.

Reproducibility and validity

A fixed seed can reproduce an inadequate simulation exactly.

Reproducibility asks whether another analyst can regenerate the computation. Adequacy asks whether the computation represents the intended linguistic process.

Joint distributions

September 23

Question: which values occurred together?

Separate one-variable distributions lose the pairing among attributes of the same outcome.

A joint distribution retains that pairing.

Several variables use one probability space

Let X:Ω→𝒳X:\Omega\rightarrow\mathcal X and Y:Ω→𝒴Y:\Omega\rightarrow\mathcal Y be random variables on the same probability space.

The ordered pair ⟨X(ω),Y(ω)⟩\langle X(\omega),Y(\omega)\rangle records two values from the same outcome ω\omega.

This is why two variables on one token do not constitute two independent observations.

Defining a joint PMF

For discrete XX and YY, define

pX,Y(x,y)≡ℙ({ω∈Ω∣X(ω)=x,Y(ω)=y}). p_{X,Y}(x,y) \equiv\mathbb{P}\!\left(\{\omega\in\Omega\mid X(\omega)=x,\ Y(\omega)=y\}\right).

After this full definition, ℙ(X=x,Y=y)\mathbb P(X=x,Y=y) abbreviates the same joint event.

Notes: joint distributions · Module 1: joint events

Declare the annotations before reading the table

Let the outcome be an annotated dialogue-turn record. Let PP code whether the turn contains a marker from a predeclared politeness-marker inventory, and let RR code whether its measured fundamental-frequency trajectory, or F0F_0 contour, meets a predeclared final-rise criterion.

These are binary annotations, not claims that politeness or intonation is intrinsically binary.

Praat: pitch analysis produces a measured pitch curve

A joint distribution of two annotations

R=1R=1 R=0R=0 total
P=1P=1 .30 .20 .50
P=0P=0 .10 .40 .50
total .40 .60 1

Reading a joint cell

The upper-left cell states

pP,R(1,1)=.30. p_{P,R}(1,1) =.30.

It assigns probability to the paired values for one dialogue turn.

The cell concerns co-occurrence within one turn. It does not by itself identify a causal direction.

Marginal PMFs do not recover pairing

Separate PMFs for XX and YY do not determine which values occur together.

Two joint tables can share the same margins while assigning different probability to the pairs.

The continuous joint distribution

A joint density fX,Yf_{X,Y} satisfies

ℙ({ω∈Ω∣⟨X(ω),Y(ω)⟩∈A})=∬AfX,Y(x,y)dxdy. \mathbb P\!\left(\{\omega\in\Omega\mid \langle X(\omega),Y(\omega)\rangle\in A\}\right) =\iint_A f_{X,Y}(x,y)\,\mathrm dx\,\mathrm dy.

Probability is area or volume under the joint density over the selected region.

Question: what remains if one coordinate is ignored?

To recover the distribution of PP alone, add across every possible value of RR.

This operation is marginalization.

What is a partition?

A partition splits an event into disjoint pieces whose union recovers the event.

For marginalization, the pieces hold one value fixed and let the value being removed vary across its codomain.

Marginalization is a partition argument

For fixed pp, the following value events partition the event for P=pP=p as rr ranges over cod⁡(R)\operatorname{cod}(R):

{ω∈Ω∣P(ω)=p}=⨄r∈cod⁡(R){ω∈Ω∣P(ω)=p,R(ω)=r}. \{\omega\in\Omega\mid P(\omega)=p\} =\biguplus_{r\in\operatorname{cod}(R)} \{\omega\in\Omega\mid P(\omega)=p,\ R(\omega)=r\}.

Countable additivity gives

ℙ(P=p)=∑r∈cod⁡(R)ℙ(P=p,R=r). \mathbb{P}(P=p) =\sum_{r\in\operatorname{cod}(R)}\mathbb{P}(P=p,R=r).

Applying the PMF definitions to both sides yields

pP(p)=∑r∈cod⁡(R)pP,R(p,r). p_P(p) =\sum_{r\in\operatorname{cod}(R)} p_{P,R}(p,r).

Computing a marginal PMF

For turns with a politeness marker,

pP(1)=pP,R(1,1)+pP,R(1,0)=.30+.20=.50. \begin{aligned} p_P(1) &=p_{P,R}(1,1)+p_{P,R}(1,0)\\ &=.30+.20\\ &=.50. \end{aligned}

We sum out pitch rise while holding politeness fixed.

Notes: marginal distributions

What marginalization removes

The marginal PMFs pPp_P and pRp_R retain the row and column totals. They do not retain how probability is paired within the four cells.

Different joint PMFs can have the same marginals.

Marginalization preserves one coordinate’s probabilities and discards its pairing with the other coordinate.

The continuous version

For two continuous measurements, probability is volume under a surface rather than mass in separate cells.

Defining a joint density

Absolutely continuous XX and YY have joint density fX,Yf_{X,Y} when, for every measurable region A⊆ℝ2A\subseteq\mathbb R^2,

ℙ(⟨X,Y⟩∈A)=∬AfX,Y(x,y)dxdy. \mathbb P\bigl(\langle X,Y\rangle\in A\bigr) =\iint_A f_{X,Y}(x,y)\,\mathrm dx\,\mathrm dy.

The density is nonnegative and integrates to one over ℝ2\mathbb R^2.

Marginal density

If XX and YY have joint density fX,Yf_{X,Y}, define the marginal density of XX by

fX(x)≡∫−∞∞fX,Y(x,y)dy. f_X(x) \equiv\int_{-\infty}^{\infty}f_{X,Y}(x,y)\,\mathrm{d}y.

The integral removes the YY coordinate.

Question: what changes after learning Y=yY=y?

Conditioning restricts attention to outcomes with a specified value and rescales their probabilities.

This is the random-variable version of conditional probability from Module 1.

Defining a conditional PMF

For any yy with pY(y)>0p_Y(y)>0, define

pX∣Y(x∣y)≡ℙ(X=x∣Y=y)=pX,Y(x,y)pY(y). p_{X\mid Y}(x\mid y) \equiv\mathbb{P}(X=x\mid Y=y) =\frac{p_{X,Y}(x,y)}{p_Y(y)}.

Module 1: conditioning as rescaling · Notes: conditional distributions

A conditional-PMF calculation

In the politeness and pitch-rise table,

pP∣R(1∣1)=pP,R(1,1)pR(1)=.30.40=.75. \begin{aligned} p_{P\mid R}(1\mid1) &=\frac{p_{P,R}(1,1)}{p_R(1)}\\ &=\frac{.30}{.40}\\ &=.75. \end{aligned}

The marginal probability is pP(1)=.50p_P(1)=.50. Conditioning on the final-rise annotation changes the probability of the politeness-marker annotation in this constructed distribution.

.75≠.50.75\neq.50, so the two annotations are not independent under this table.

Deriving the chain rule for random variables

Multiply the conditional-PMF definition by pY(y)p_Y(y):

pX,Y(x,y)=pX∣Y(x∣y)pY(y). p_{X,Y}(x,y) =p_{X\mid Y}(x\mid y)p_Y(y).

This is the two-variable chain rule.

It says that a joint cell can be constructed from one marginal probability and one conditional probability.

Defining a conditional density

For values yy with fY(y)>0f_Y(y)>0, define

fX∣Y(x∣y)≡fX,Y(x,y)fY(y). f_{X\mid Y}(x\mid y) \equiv\frac{f_{X,Y}(x,y)}{f_Y(y)}.

Thus fX,Y(x,y)=fX∣Y(x∣y)fY(y)f_{X,Y}(x,y)=f_{X\mid Y}(x\mid y)f_Y(y) wherever the conditional density is defined.

Question: when does conditioning make no difference?

If learning Y=yY=y leaves the complete distribution of XX unchanged, then XX and YY are independent.

Independence as conditional invariance

For every yy with pY(y)>0p_Y(y)>0, independence requires

pX∣Y(x∣y)=pX(x) p_{X\mid Y}(x\mid y)=p_X(x)

for every xx. Learning Y=yY=y does not change the distribution of XX.

Module 1: independence

Deriving the independence factorization

A factorization rewrites a joint distribution as a product. Substitute conditional invariance into the chain rule:

pX,Y(x,y)=pX∣Y(x∣y)pY(y)=pX(x)pY(y). \begin{aligned} p_{X,Y}(x,y) &=p_{X\mid Y}(x\mid y)p_Y(y)\\ &=p_X(x)p_Y(y). \end{aligned}

The product definition of independence

For discrete XX and YY, define

X⟂⟂Y≡[pX,Y(x,y)=pX(x)pY(y) for every x,y]. X\perp\!\!\!\perp Y \quad\equiv\quad \bigl[p_{X,Y}(x,y)=p_X(x)p_Y(y) \text{ for every }x,y\bigr].

For absolutely continuous variables, replace the PMFs by densities; the equality may fail only on a set with zero area under integration.

Independence is a claim about the entire joint distribution, not just zero correlation.

Can dependence disappear within groups?

Two variables may be associated after pooling observations but independent within every value of a third variable ZZ.

This is conditional independence. We must state the conditioning variable.

Conditional invariance given ZZ

For values with pY,Z(y,z)>0p_{Y,Z}(y,z)>0, conditional independence requires

pX∣Y,Z(x∣y,z)=pX∣Z(x∣z). p_{X\mid Y,Z}(x\mid y,z)=p_{X\mid Z}(x\mid z).

This is the conditional-invariance characterization of X⟂⟂Y∣ZX\perp\!\!\!\perp Y\mid Z.

Conditional independence factorization

For every zz with pZ(z)>0p_Z(z)>0, define

X⟂⟂Y∣Z≡[pX,Y∣Z(x,y∣z)=pX∣Z(x∣z)pY∣Z(y∣z) for every x,y]. X\perp\!\!\!\perp Y\mid Z \quad\equiv\quad \left[ p_{X,Y\mid Z}(x,y\mid z) =p_{X\mid Z}(x\mid z)p_{Y\mid Z}(y\mid z) \text{ for every }x,y \right].

The conditional chain rule makes this definition equivalent to conditional invariance wherever the latter is defined.

Notes: conditional independence

X⟂⟂Y∣ZX\perp\!\!\!\perp Y\mid Z does not imply X⟂⟂YX\perp\!\!\!\perp Y after pooling over ZZ.

Linguistic instance: hold clause type fixed

Let OO mark subject omission, VV mark verb-final order, and CC mark main versus subordinate clause type in a constructed corpus model.

The claim

O⟂⟂V∣C O\perp\!\!\!\perp V\mid C

says that within each clause type, learning VV does not change the distribution of OO. It does not say that omission and order are independent in the pooled corpus.

Within-group factorization

For main clauses,

pO,V∣C(1,1∣main)=.80(.75)=.60. p_{O,V\mid C}(1,1\mid\mathrm{main})=.80(.75)=.60.

For subordinate clauses, the corresponding value is .20(.25)=.05.20(.25)=.05.

Pooling can restore dependence

If both clause types have probability .50.50,

pO,V(1,1)=.60(.50)+.05(.50)=.325, p_{O,V}(1,1)=.60(.50)+.05(.50)=.325,

while pO(1)pV(1)=.50(.50)=.25p_O(1)p_V(1)=.50(.50)=.25. Conditional independence does not imply pooled independence.

Question: what is the mean within one group?

Conditional expectations average a response using probabilities within one declared group.

Defining conditional expectation

For discrete XX and any yy with pY(y)>0p_Y(y)>0, define

𝔼[X∣Y=y]≡∑x∈cod⁡(X)xpX∣Y(x∣y). \mathbb E[X\mid Y=y] \equiv\sum_{x\in\operatorname{cod}(X)} x\,p_{X\mid Y}(x\mid y).

This is the mean of the conditional distribution at Y=yY=y.

Notes: conditional expectation

Reading-time means by predictability

reading time low high
400 ms .20.20 .60.60
500 ms .50.50 .30.30
600 ms .30.30 .10.10

𝔼[Y∣X=low]=510 ms,𝔼[Y∣X=high]=450 ms. \mathbb E[Y\mid X=\mathrm{low}]=510\text{ ms}, \qquad \mathbb E[Y\mid X=\mathrm{high}]=450\text{ ms}.

The marginal mean mixes conditions

If high predictability has probability .70.70,

𝔼[Y]=.70(450)+.30(510)=468 ms. \mathbb E[Y]=.70(450)+.30(510)=468\text{ ms}.

Changing the condition mixture can change the marginal mean without changing either conditional distribution.

Begin with the joint distribution

For integrable discrete variables,

𝔼[X]=∑x∈cod⁡(X)x∑y∈cod⁡(Y)pX,Y(x,y)=∑y∈cod⁡(Y)∑x∈cod⁡(X)xpX,Y(x,y). \begin{aligned} \mathbb E[X] &=\sum_{x\in\operatorname{cod}(X)}x \sum_{y\in\operatorname{cod}(Y)}p_{X,Y}(x,y)\\ &=\sum_{y\in\operatorname{cod}(Y)} \sum_{x\in\operatorname{cod}(X)} x\,p_{X,Y}(x,y). \end{aligned}

The first line marginalizes over YY. The second line changes the order of the two sums.

Refactor each joint probability

Use

pX,Y(x,y)=pX∣Y(x∣y)pY(y). p_{X,Y}(x,y)=p_{X\mid Y}(x\mid y)p_Y(y).

Then

𝔼[X]=∑y∈cod⁡(Y)[∑x∈cod⁡(X)xpX∣Y(x∣y)]pY(y). \mathbb E[X] =\sum_{y\in\operatorname{cod}(Y)} \left[ \sum_{x\in\operatorname{cod}(X)} x\,p_{X\mid Y}(x\mid y) \right]p_Y(y).

The bracketed sum is the conditional mean at Y=yY=y.

The law of total expectation

Substituting the conditional-mean definition gives

𝔼[X]=∑y∈cod⁡(Y)𝔼[X∣Y=y]pY(y). \mathbb E[X] =\sum_{y\in\operatorname{cod}(Y)} \mathbb E[X\mid Y=y]p_Y(y).

The marginal mean is a probability-weighted average of the conditional means.

From contextual probability to information

A language model assigns a probability to each candidate next-word value given the preceding context. The observed token’s surprisal uses the probability assigned to its realized value: lower probability means greater surprisal.

The unit is a bit when logarithms use base 2. Halving the contextual probability adds one bit.

Question: do paired values vary together?

Covariance averages the product of two paired deviations. Pairing within the same outcome is essential.

For a constructed illustration, let WW be the next-word value and wobsw_{\mathrm{obs}} its observed realization. Define that token’s contextual surprisal as

X≡−log⁡2ℙ(W=wobs∣preceding context), X\equiv-\log_2\mathbb P(W=w_{\mathrm{obs}}\mid\text{preceding context}),

and let YY be its reading time in milliseconds.

Defining covariance

For variables with finite second moments, define

Cov⁡(X,Y)≡𝔼[(X−𝔼[X])(Y−𝔼[Y])]. \operatorname{Cov}(X,Y) \equiv\mathbb E\!\left[(X-\mathbb E[X])(Y-\mathbb E[Y])\right].

Its units are the product of the units of XX and YY.

Notes: covariance

Work a covariance

Give four token outcomes equal mass and paired values

⟨X,Y⟩∈{⟨1,200⟩,⟨1,250⟩,⟨3,350⟩,⟨3,400⟩}. \langle X,Y\rangle \in\{\langle1,200\rangle,\langle1,250\rangle, \langle3,350\rangle,\langle3,400\rangle\}.

Then 𝔼[X]=2\mathbb E[X]=2 bits, 𝔼[Y]=300\mathbb E[Y]=300 ms, and

Cov⁡(X,Y)=75 bit ms. \operatorname{Cov}(X,Y)=75\text{ bit ms}.

The positive sign describes the pairing; it does not explain why less predictable words took longer.

Smith and Levy on predictability and reading time

Deriving the covariance identity

Expanding the product and applying linearity gives

Cov⁡(X,Y)=𝔼[XY−X𝔼[Y]−𝔼[X]Y+𝔼[X]𝔼[Y]]=𝔼[XY]−𝔼[X]𝔼[Y]−𝔼[X]𝔼[Y]+𝔼[X]𝔼[Y]=𝔼[XY]−𝔼[X]𝔼[Y]. \begin{aligned} \operatorname{Cov}(X,Y) &=\mathbb E[XY-X\mathbb E[Y]-\mathbb E[X]Y +\mathbb E[X]\mathbb E[Y]]\\ &=\mathbb E[XY]-\mathbb E[X]\mathbb E[Y] -\mathbb E[X]\mathbb E[Y] +\mathbb E[X]\mathbb E[Y]\\ &=\mathbb E[XY]-\mathbb E[X]\mathbb E[Y]. \end{aligned}

This identity does not define covariance; it follows from the preceding definition.

Pairing is essential

Reversing the reading-time order leaves the two marginal distributions unchanged but changes which values occur together.

Covariance belongs to the joint distribution. The marginals do not identify it.

Covariance retains measurement units

If surprisal is measured in bits and reading time in milliseconds, then covariance is measured in bit milliseconds.

Changing the measurement scale changes the numerical covariance.

Zero covariance does not imply independence

Covariance summarizes linear co-movement through paired products. A curved or balanced dependence can have covariance zero.

Independence implies zero covariance when the relevant expectations exist, but the reverse does not generally hold.

How can covariance be made unitless?

Divide covariance by the two standard deviations. The result is correlation.

Correlation

When both standard deviations are positive, define

ρX,Y≡Cov⁡(X,Y)σXσY. \rho_{X,Y} \equiv \frac{\operatorname{Cov}(X,Y)} {\sigma_X\sigma_Y}.

Correlation is unitless and satisfies −1≤ρX,Y≤1-1\leq\rho_{X,Y}\leq1.

Notes: correlation

Work a linguistic correlation

The English dative alternation contrasts the double-object form gave her the book with the prepositional form gave the book to her.

Let X=1X=1 mark a recipient expressed as a pronoun and Y=1Y=1 mark the double-object construction in a constructed table.

If 𝔼[X]=.50\mathbb E[X]=.50, 𝔼[Y]=.60\mathbb E[Y]=.60, and 𝔼[XY]=.40\mathbb E[XY]=.40, then

Cov⁡(X,Y)=.10,ρX,Y=.10.50.24≈.408. \operatorname{Cov}(X,Y)=.10, \qquad \rho_{X,Y}=\frac{.10}{.50\sqrt{.24}}\approx.408.

Bresnan et al. on predictors of the dative alternation

Correlation has a limited interpretation

Correlation summarizes linear association under the joint distribution. It does not determine the complete dependence structure and does not identify a causal direction.

Correlation answers a narrow descriptive question. It does not explain why the association arose.

Module summary

  1. Random variables map outcomes to values.
  2. PMFs assign mass, PDFs assign density, and CDFs accumulate probability.
  3. Expectations summarize specified functions when they exist.
  4. Distribution families pair support with probability form and process assumptions.
  5. Joint distributions determine marginals, conditionals, and paired summaries.

Which associations survive after a relevant linguistic grouping variable is held fixed?

Return to the module notes · Review the dependency graph