Statistical Methods in Linguistics
University of Rochester
September 14, 16, 21 and 23, 2026
A contextualized pronoun token has a spelling, an annotated case value, a speaker, and a duration. Which part of that record should enter a particular analysis?
Module 1 assigned probability to events, that is, observable sets of complete outcomes.
But a linguistic question usually targets one attribute of each outcome:
If an analysis concerns duration, then replacing each complete token by its duration should preserve the events needed for duration questions.
But it should not erase the token record that supports other questions.
A random variable extracts one value from each outcome. Its preimages connect value questions back to the event space from Module 1.
Each meeting starts from a linguistic question, defines the probability object that answers it, and checks the definition on a linguistic example.
September 14
Suppose the random experiment selects one complete vowel token.
| outcome | word | speaker | duration |
|---|---|---|---|
| heed | s01 | 112 ms | |
| hid | s01 | 83 ms | |
| heed | s02 | 126 ms |
The outcome retains the word, speaker, waveform, vowel label, and duration.
Let . Define
Thus , , and .
The event that duration exceeds 100 ms is
The event contains vowel-token outcomes, not the bare values 112 and 126.
Define
The same outcome may satisfy and .
The operation selects duration values but returns the corresponding outcome identifiers.
Keeping only 112, 83, and 126 removes speaker and word identity.
A later analysis could no longer ask whether durations repeat within speakers or differ across words.
A random variable maps complete outcomes to values. Retaining the outcome identifiers keeps each extracted value connected to the observation process.
Basically, a random variable is a rule that reads one value from each outcome.
Write the outcome-to-value mapping as
The domain contains outcomes. The codomain contains possible values. We add the formal measurability requirement after examining preimages.
For a random variable , write its declared value space as
Thus a sum over possible values uses . The support may be smaller.
For a discrete variable, its support is the subset of codomain values with positive probability mass.
For each token , define
Thus , , and .
The numbers are measured values in milliseconds. They are not the speech tokens themselves.
Suppose we ask whether a selected token lasts more than 100 ms.
Prediction: the answer should be a set of token outcomes, since Module 1 assigns probability to events rather than to ungrounded values.
For a set of values , define the shorthand
This set contains outcomes, not bare values.
The preimage gathers exactly the outcomes whose values land in .
For the three-outcome illustration, let . Then
This is the duration-threshold event in .
A value-space question becomes probabilistically meaningful by pointing back to an event in .
We need every value set we might ask about to point back to an observable event.
This requirement is measurability.
Let be the collection of value sets we permit in probability questions.
Like the event space from Module 1, must also permit:
The structure is a measurable space.
The function is measurable when
Thus every permitted value question has an observable preimage event.
A random variable is a function satisfying this condition.
Notes: measurable random variables · Module 1: sigma-algebras
Let each outcome be a contextualized pronoun token. Define as its combined form-and-case label, with codomain
An integer may store a position in this ordered codomain, but that code has no linguistic meaning apart from the declared correspondence.
For the first three labels in the declared order,
More generally,
The measurability condition requires this preimage to belong to for every .
The set on the right contains token records, not three abstract pronoun spellings.
The probability space remains on . The random variable induces probabilities for measurable sets of values.
Prediction: the probability of a value set must equal the probability of its preimage event.
The distribution of assigns each the probability
As a function of , this assignment is a probability measure on . We retain for the probability measure on and introduce no second measure symbol.
This question identifies discrete distributions. We ask separately whether a non-discrete distribution has a density.
A random variable is discrete when some finite or countably infinite set satisfies .
Annotated grammatical person has finite categorical support. A token count or an utterance-length count has finite or countably infinite numerical support.
Discrete does not mean numerical or unordered. It means that a listable set carries all probability.
The codomain lists values the function may return. The support contains the values with positive probability:
is finite. An utterance-length support such as is countably infinite. Both are discrete.
For observed lengths
the distinct observed values are . The modeled support may still contain , , , and larger integers.
Grammatical person is discrete without a numerical distance between labels. Utterance length is discrete and ordered, with interpretable differences.
Discreteness concerns whether the probability-carrying values are separate and listable.
For a discrete variable, separate codomain values can receive probability mass.
The function that records those probabilities is the probability mass function (PMF).
For a discrete random variable , its probability mass function (PMF) is defined by
The subscript names the random variable; the argument is a possible value.
Its support is the subset of codomain values with positive mass:
Its PMF satisfies
For any ,
Suppose counts syllables in a sampled word token.
| syllable count | 1 | 2 | 3 | 4 |
|---|---|---|---|---|
| .42 | .33 | .17 | .08 |
Because and are disjoint,
For a discrete variable, an event probability is the sum of the masses at the codomain values in that event. Zero-mass values contribute zero.
If , then
under the stated model. This does not imply that seven-syllable words are impossible in every linguistic population.
Dropping the row with mass leaves total mass , which is not a PMF.
Renormalizing the remaining rows defines a different model, conditional on at most three syllables.
A duration is recorded at finite precision, but that storage convention does not force the model to have discrete support.
Prediction: interval probabilities should come from area, while individual points receive zero mass.
A Borel set belongs to the smallest collection of subsets of that contains every open interval and is closed under complements and countable unions.
Thus ordinary intervals, threshold sets, and sets formed from them by these operations are Borel.
A real-valued random variable is absolutely continuous when there is a function such that, for every Borel set ,
The function is a probability density function.
Recall from Module 1 that a formant-frequency estimate records the estimated center frequency of a vocal-tract resonance. Suppose returns the first two such estimates from a vowel token. Let the idealized value space be and define
The codomain is uncountable. That fact alone does not make the distribution absolutely continuous.
Absolute continuity requires probabilities to be recovered by integrating a density over ordinary two-dimensional area.
The units are hertz; the pair is an acoustic measurement, not a phoneme or vowel category.
Under a continuous model,
while an interval around that point may receive positive probability:
A stored value of 18 ms may be a rounded observation of an interval-valued quantity, or it may be the discrete output of the measurement process.
The scientific target determines which representation is appropriate.
| question | discrete | continuous |
|---|---|---|
| values | separate and listable | fill a region |
| one value | may have positive mass | has probability zero |
| event probability | add masses | integrate density |
The density of satisfies
Conversely, any measurable nonnegative function with integral one defines an absolutely continuous probability distribution by integration.
Suppose
Then
This is a model for a measured token duration. It is not a claim that durations are actually uniform.
If is measured in syllables per second, then is measured in seconds per syllable.
Interval width times density is unitless probability.
Height 2 across an interval of width has total area
Density height is not bounded by one. Probability is.
For an absolutely continuous random variable,
Thus does not imply .
Density height is not point probability. Only an integral of density gives probability.
Both a PMF and a PDF answer probability questions, but they do so differently.
The cumulative distribution function (CDF) instead asks one shared threshold question: how much probability lies at or below ?
Every real-valued random variable has a cumulative distribution function (CDF), defined by
The same definition applies to discrete, continuous, and mixed distributions.
For a discrete ,
For an absolutely continuous ,
These are consequences of the CDF definition.
PMFs, PDFs, and CDFs are different representations of distributions, not interchangeable names.
A PMF assigns mass to values. A PDF assigns density whose integral is probability. A CDF gives the probability at or below a threshold.
September 16
A plain average gives equal weight to observed rows. A distributional average weights every possible value by its probability.
This probability-weighted average is the expected value.
A variable is integrable when its expected absolute value is finite:
This requirement prevents undefined cancellation between infinite positive and negative contributions.
For an integrable random variable , define
Suppose counts syllables in a sampled word token and takes values , , and with masses , , and . Then
An expectation need not be a possible syllable count for any single token.
describes the distribution, not an individual word token.
Some questions concern a transformed value, such as squared distance from a reference point.
Prediction: transform each possible value first, then average those transformed values using the original probabilities.
When is integrable, define
in the discrete case, with the analogous integral in the continuous case.
In general, .
For with masses and ,
Thus . The function encodes the property being averaged.
For a linear transformation , averaging then transforming gives the same result as transforming then averaging.
Let us derive that claim before using it.
If marks a target event and , then
For probabilities , , and ,
For constants and ,
If each is integrable,
This equality does not require independence. Dependence can change the distribution and variance of the sum.
Notes: linearity of expectation
Expectation adds across variables even when their values co-vary.
The indicator values may be dependent. Linearity still gives
The corresponding product rule does require additional conditions:
in general.
No. The positive and negative parts of the weighted integral must not both diverge.
For with ,
The positive and negative parts of are both infinite. Thus is undefined.
Integrating over gives zero for every . The limit of these symmetric truncations is the symmetric principal value, not an expectation.
Every finite sample of Cauchy draws has a finite arithmetic mean. An extreme draw can still move that mean by an arbitrarily large amount as the sample grows.
The model supplies no finite population mean to which the sample means must converge.
Signed deviations cancel. Squaring them yields nonnegative distances and weights large deviations more heavily.
Their expectation is the variance.
Let equal 2 with probability one. Let equal 1 or 3, each with probability .
The second moment is .
It is finite when
This condition ensures that the squared deviations used by variance have a finite average.
For a random variable with a finite second moment, define
Variance is an expected squared deviation from the mean.
Writing ,
Variance squares the measurement unit. Taking its square root returns to the original unit.
The standard deviation is defined by
If is measured in milliseconds, has units ms and has units ms.
If
then
Standard deviation is the root mean squared distance from the mean, not the average absolute distance.
Equal means and standard deviations do not guarantee equal tails, symmetry, or modality.
The mean and standard deviation do not determine how much probability lies within .
Coverage claims require an additional distribution-shape assumption.
Variance uses the second power. Higher centered powers can summarize other aspects of distribution shape, when the corresponding expectations exist.
When it exists, the th central moment is defined by
Thus and .
For with masses and ,
Odd powers retain direction. Higher powers give more weight to values far from the mean.
Different distributions can share the same first several moments.
Thus moments summarize specified centered powers; they do not determine the observation process or the full distributional shape.
An expectation summarizes a specified function of a distribution. Its meaning depends on that function, the existence of the expectation, and the measurement units.
September 21
A distribution family links a support to a particular probability form.
Thus family choice must follow the observation process and the research question, not merely the shape of a histogram.
A parameter is a fixed value that selects one distribution from a family. Changing it may change location, spread, shape, or several of these properties.
The statement
means that the distribution of belongs to family with parameter . The symbol does not mean numerical equality.
It says that has a distribution in the named family under the stated parameter convention.
Suppose one contextualized pronoun token receives exactly one form-and-case label from a finite list.
The categorical distribution assigns one probability to each label.
For labels and with ,
means
Order the pronoun labels as shown in the graph and define the probability vector
The th entry is the probability of the th form-and-case label. Reordering the labels without applying the same permutation to changes the distribution.
The labels encode two annotated attributes. They are not fourteen distinct English spellings.
Case=Acc?The categorical record retains all fourteen labels. If the question concerns one binary contrast, we can map those labels to 1 and 0.
Prediction: different spellings may receive the same value, and a syncretic spelling may receive different values in different contexts.
Define the accusative event
Define the accusative-case indicator
The declared codomain is even though the sample space contains fourteen pronoun outcomes. For ,
means
Using ,
The parameter is the probability of annotated accusative case because the definition of assigns Case=Acc outcomes the value .
Bernoulli parameters inherit their meaning from the event coded as 1.
For ,
The mean equals the probability of the event coded as 1.
If the coding is reversed, the parameter becomes the probability of the complementary event.
Thus is not an intrinsic property of the spelling. It inherits its meaning from the declared outcome-to-value mapping.
Suppose we predeclare token opportunities and count how many have Case=Acc.
Prediction: a binomial model needs a fixed , independent indicators, and one common success probability .
Let be independent Bernoulli variables with common parameter , and define . Then
with support .
The binomial coefficient counts the ways to choose positions from . Each sequence with ones has probability . Thus
counts the distinct indicator sequences that produce the same total .
For , , and ,
The calculation concerns exactly two target outcomes among six declared opportunities.
For ,
The model requires a fixed number of opportunities, independent indicators, and one common success probability.
The count does not retain which two positions were successes, and the same proportion can arise from different values of .
If selecting one clause changes which clauses remain available, later draws are not independent Bernoulli trials.
The hypergeometric distribution represents a count from a fixed finite population sampled without replacement.
Suppose an archive contains clauses, of which are passive. Select clauses without replacement. If counts passives, then
The support satisfies
One unit is a clause record, and passive voice is a declared annotation on that record.
Under English Universal Dependencies, Voice=Pass marks a verbal past participle used as the predicate of a passive clause.
Notes: hypergeometric distributions · UD English Voice · R manual
For , , , and ,
The factor reduces variance because sampling one clause changes what remains.
A binomial count from independent draws with probability also has mean , but variance .
The hypergeometric variance is smaller because the draws are dependent after clauses are removed.
Let each opportunity be a Bernoulli trial, and let count failures before the first success.
Before naming the family, we need probabilities over that sum to one.
A geometric series multiplies each successive term by the same ratio. Begin with
Thus
is a PMF with countably infinite support.
Let count failures before the first success in independent Bernoulli trials with common success probability . Define
by the PMF
This is the convention used by dgeom in R. The total-trial count is .
Notes: geometric distributions · R manual: geometric distribution
For ,
The first quantity fixes the first target after three failures; the second requires four initial failures.
Let . Differentiating the geometric series gives
Thus
The mode is a value with the greatest probability mass. A mode at the smallest possible value is a boundary mode.
For every ,
Thus . If a positive count is , every geometric model has its mode at .
This is a model prediction. We can now compare it with an empirical distribution.
The negative binomial distribution counts failures before the th success. The additional parameter permits an interior mode.
This is a family comparison, not yet a claim that dictionary symbols arise from literal Bernoulli trials.
Let count failures before the th success, where . Define
by
When , the binomial coefficient equals one and this PMF is exactly geometric.
With fixed, increasing increases
Set so that the observed support begins at zero. Under this negative-binomial convention,
For a count distribution, dispersion compares variance with the mean.
Writing gives
for every finite .
Thus every finite negative-binomial model is overdispersed relative to its mean.
For , , and ,
The final draw is the third target; the four failures occupy the first six positions.
A model PMF assigns probabilities to possible values.
An empirical distribution assigns each observed value its relative frequency.
Use every noncomment pronunciation entry in CMU Pronouncing Dictionary 0.7b. For entry , define
The unit is a pronunciation entry, not a corpus token or necessarily a unique spelling type.
CMUdict file-format documentation · Notes: the worked example
The empirical relative-frequency function has its mode at six symbols. A fitted geometric PMF must have its mode at one symbol.
Hold the observed data fixed. The likelihood is the joint PMF evaluated at observed discrete data, or the joint density evaluated at observed continuous data, treated as a function of the parameters.
A maximum-likelihood estimate is a parameter value that maximizes this function, when such a value exists. A larger likelihood means a better fit to these observations within the compared family.
Prediction: a finite cannot match underdispersed data, whose variance is below their mean.
For the pronunciation-entry lengths, the mean of is 5.381 and its variance is 4.566.
For each , estimate and evaluate the likelihood of all observed entry lengths.
The likelihood increases toward , the Poisson boundary derived next.
The extra parameter fixes the geometric mode restriction, but every finite negative-binomial model remains overdispersed relative to its mean.
The negative-binomial generalization removes the geometric family’s boundary-mode restriction. For these dictionary entries, the maximum-likelihood fit approaches its Poisson limit rather than a finite .
An added parameter can repair one mismatch while leaving another. Model criticism must name both.
The fitted negative-binomial sequence approaches a Poisson distribution while keeping its mean fixed.
Let us state that limit before interpreting a Poisson count.
Fix and let
Then, for every ,
Thus the negative-binomial family approaches as while its mean remains .
A count has no interpretable rate without an exposure: tokens, clauses, speakers, minutes, or another declared unit.
If events occur at rate per exposure unit, a Poisson model for exposure specifies
with
For ,
The first probability concerns no filled pauses; the second concerns exactly two in the declared exposure.
The count model implies
A Poisson process with a constant rate additionally specifies , independent counts in disjoint intervals, and the same expected count for intervals of the same length.
A Poisson marginal distribution does not by itself license a story about a Poisson process.
On a finite interval, constant density produces probability proportional to interval length.
This is the continuous uniform distribution.
For ,
means
For ,
Equal density height means equal probability for intervals of equal length, not equal positive probability at every point.
Let be a computer-generated draw used to assign a trial to one of two lists:
Here uniformity describes the randomization mechanism. It does not assert that an acoustic measurement is naturally uniform.
Proportions and bounded continuous responses may cluster near the center, near one boundary, or near both boundaries.
The beta family can represent these contrasting shapes through two positive parameters.
For , the beta function supplies the normalizing area
Then
has possible values in and density
If and
then, for ,
The change-of-scale factor is required for the transformed density to integrate to one.
For ,
Holding this ratio fixed while increasing concentrates the distribution around the same mean.
controls the mean, while controls concentration when the mean is held fixed.
Suppose is a participant’s sentence-acceptability rating, recorded with a continuous slider and rescaled to .
A beta model treats as a measurement, not as a grammatical category. If the task permits exact endpoint responses 0 and 1, an ordinary beta distribution is not sufficient without an endpoint-handling component.
The normal family uses one parameter for its center and another for its spread over the real line.
This support assumption matters: an untransformed duration or proportion cannot literally take every real value.
For ,
means
Standardization expresses a value’s distance from the model mean in standard-deviation units.
Define
If , then . Standardization removes the physical unit and retains relative position within this family.
A standardized value reports position in standard-deviation units under the assumed normal model.
For ,
Thus under this normal model.
Let be a token duration and . A normal model for states
This is a distributional assumption about a transformed acoustic measurement. It is not implied by the fact that a duration histogram looks roughly bell-shaped.
Squaring standard-normal variables makes every contribution nonnegative. Adding independent squared values yields a chi-squared variable.
The count is called its degrees of freedom: it records the number of independent squared contributions.
For , let be independent standard normal variables and define
Then , with and .
Suppose are independent standardized discrepancies for three predeclared acoustic measurements. Then
under the standard-normal assumptions. The degrees of freedom count independent squared contributions, not speech tokens.
For , , and ,
Under the independent standard-normal construction, and .
For ,
Degrees of freedom determine both center and spread.
For the constructed value with ,
This is probability under the reference distribution, not the probability that a hypothesis is true.
In an by table with fixed row and column totals,
The symbol counts independent pieces of variation remaining after constraints, not necessarily observations.
Dividing a standard-normal variable by an independent estimated scale adds uncertainty. The resulting Student’s t distribution has heavier tails, that is, it places more probability far from zero.
Let and be independent. Define
Then . Because can be close to zero, extreme values of are more probable than extreme values under a standard normal distribution.
For and ,
The corresponding standard-normal probability is about because has heavier tails.
An estimator is a rule that maps a sample to an estimate of an unknown parameter. A distribution may describe a standardized estimator whose scale was estimated.
It may instead be chosen as an observation model, that is, a distribution for a transformed linguistic response.
Those are different claims. Naming the random variable and its generating construction keeps them apart.
A distribution family specifies a support and a mass or density function. A process interpretation requires additional assumptions about how observations are generated.
A CDF partitions probability from zero to one. A uniform draw can select one of those probability intervals.
This is the intuition behind inverse transform sampling.
For inverse transform sampling, the generalized inverse returns the leftmost value at which the CDF reaches . Formally,
If , define .
The generalized inverse satisfies exactly when . Thus
For ordered values , define
Draw and return the first for which .
Cumulative mass turns one uniform value into one draw from the target PMF.
| label | mass | endpoint | interval |
|---|---|---|---|
| Nom | |||
| Acc | |||
| Gen |
Suppose a constructed three-label PMF assigns to (Nom, Acc, Gen). Its cumulative cutpoints are .
If , return Acc, since .
The returned value is an annotated category label, not a spelling or pronunciation.
The endpoints must be the cumulative sums .
If the first endpoint were , the first label would be simulated too rarely. Large-sample proportions would expose the error.
A pseudorandom algorithm is deterministic once its initial state is fixed.
A random seed records that initial state.
A random seed fixes the initial state of a pseudorandom number generator. Reusing the seed reproduces the same sequence of draws.
It does not validate the distribution, independence assumptions, or sampling representation.
Set the seed once before a related block of random draws. Resetting the same seed inside a loop repeats the same simulated result.
Record the seed with the code and place it before the first draw that affects the reported result.
Changing the seed provides a check on Monte Carlo variation. A conclusion that changes materially across seeds may require more draws or an implementation check.
More draws reduce simulation variation. They do not repair the wrong target distribution.
A fixed seed can reproduce an inadequate simulation exactly.
Reproducibility asks whether another analyst can regenerate the computation. Adequacy asks whether the computation represents the intended linguistic process.
September 23
Separate one-variable distributions lose the pairing among attributes of the same outcome.
A joint distribution retains that pairing.
Let and be random variables on the same probability space.
The ordered pair records two values from the same outcome .
This is why two variables on one token do not constitute two independent observations.
For discrete and , define
After this full definition, abbreviates the same joint event.
Let the outcome be an annotated dialogue-turn record. Let code whether the turn contains a marker from a predeclared politeness-marker inventory, and let code whether its measured fundamental-frequency trajectory, or contour, meets a predeclared final-rise criterion.
These are binary annotations, not claims that politeness or intonation is intrinsically binary.
| total | |||
|---|---|---|---|
| .30 | .20 | .50 | |
| .10 | .40 | .50 | |
| total | .40 | .60 | 1 |
The upper-left cell states
It assigns probability to the paired values for one dialogue turn.
The cell concerns co-occurrence within one turn. It does not by itself identify a causal direction.
Separate PMFs for and do not determine which values occur together.
Two joint tables can share the same margins while assigning different probability to the pairs.
A joint density satisfies
Probability is area or volume under the joint density over the selected region.
To recover the distribution of alone, add across every possible value of .
This operation is marginalization.
A partition splits an event into disjoint pieces whose union recovers the event.
For marginalization, the pieces hold one value fixed and let the value being removed vary across its codomain.
For fixed , the following value events partition the event for as ranges over :
Countable additivity gives
Applying the PMF definitions to both sides yields
For turns with a politeness marker,
We sum out pitch rise while holding politeness fixed.
The marginal PMFs and retain the row and column totals. They do not retain how probability is paired within the four cells.
Different joint PMFs can have the same marginals.
Marginalization preserves one coordinate’s probabilities and discards its pairing with the other coordinate.
For two continuous measurements, probability is volume under a surface rather than mass in separate cells.
Absolutely continuous and have joint density when, for every measurable region ,
The density is nonnegative and integrates to one over .
If and have joint density , define the marginal density of by
The integral removes the coordinate.
Conditioning restricts attention to outcomes with a specified value and rescales their probabilities.
This is the random-variable version of conditional probability from Module 1.
For any with , define
Module 1: conditioning as rescaling · Notes: conditional distributions
In the politeness and pitch-rise table,
The marginal probability is . Conditioning on the final-rise annotation changes the probability of the politeness-marker annotation in this constructed distribution.
, so the two annotations are not independent under this table.
Multiply the conditional-PMF definition by :
This is the two-variable chain rule.
It says that a joint cell can be constructed from one marginal probability and one conditional probability.
For values with , define
Thus wherever the conditional density is defined.
If learning leaves the complete distribution of unchanged, then and are independent.
For every with , independence requires
for every . Learning does not change the distribution of .
A factorization rewrites a joint distribution as a product. Substitute conditional invariance into the chain rule:
For discrete and , define
For absolutely continuous variables, replace the PMFs by densities; the equality may fail only on a set with zero area under integration.
Independence is a claim about the entire joint distribution, not just zero correlation.
Two variables may be associated after pooling observations but independent within every value of a third variable .
This is conditional independence. We must state the conditioning variable.
For values with , conditional independence requires
This is the conditional-invariance characterization of .
For every with , define
The conditional chain rule makes this definition equivalent to conditional invariance wherever the latter is defined.
Notes: conditional independence
does not imply after pooling over .
Let mark subject omission, mark verb-final order, and mark main versus subordinate clause type in a constructed corpus model.
The claim
says that within each clause type, learning does not change the distribution of . It does not say that omission and order are independent in the pooled corpus.
For main clauses,
For subordinate clauses, the corresponding value is .
If both clause types have probability ,
while . Conditional independence does not imply pooled independence.
Conditional expectations average a response using probabilities within one declared group.
For discrete and any with , define
This is the mean of the conditional distribution at .
| reading time | low | high |
|---|---|---|
| 400 ms | ||
| 500 ms | ||
| 600 ms |
If high predictability has probability ,
Changing the condition mixture can change the marginal mean without changing either conditional distribution.
For integrable discrete variables,
The first line marginalizes over . The second line changes the order of the two sums.
Use
Then
The bracketed sum is the conditional mean at .
Substituting the conditional-mean definition gives
The marginal mean is a probability-weighted average of the conditional means.
A language model assigns a probability to each candidate next-word value given the preceding context. The observed token’s surprisal uses the probability assigned to its realized value: lower probability means greater surprisal.
The unit is a bit when logarithms use base 2. Halving the contextual probability adds one bit.
Covariance averages the product of two paired deviations. Pairing within the same outcome is essential.
For a constructed illustration, let be the next-word value and its observed realization. Define that token’s contextual surprisal as
and let be its reading time in milliseconds.
For variables with finite second moments, define
Its units are the product of the units of and .
Give four token outcomes equal mass and paired values
Then bits, ms, and
The positive sign describes the pairing; it does not explain why less predictable words took longer.
Expanding the product and applying linearity gives
This identity does not define covariance; it follows from the preceding definition.
Reversing the reading-time order leaves the two marginal distributions unchanged but changes which values occur together.
Covariance belongs to the joint distribution. The marginals do not identify it.
If surprisal is measured in bits and reading time in milliseconds, then covariance is measured in bit milliseconds.
Changing the measurement scale changes the numerical covariance.
Covariance summarizes linear co-movement through paired products. A curved or balanced dependence can have covariance zero.
Independence implies zero covariance when the relevant expectations exist, but the reverse does not generally hold.
Divide covariance by the two standard deviations. The result is correlation.
When both standard deviations are positive, define
Correlation is unitless and satisfies .
The English dative alternation contrasts the double-object form gave her the book with the prepositional form gave the book to her.
Let mark a recipient expressed as a pronoun and mark the double-object construction in a constructed table.
If , , and , then
Correlation summarizes linear association under the joint distribution. It does not determine the complete dependence structure and does not identify a causal direction.
Correlation answers a narrow descriptive question. It does not explain why the association arose.
Which associations survive after a relevant linguistic grouping variable is held fixed?