Statistical Methods in Linguistics
University of Rochester
September 28 and 30, October 5, 2026
What can observed speakers, tokens, items, or measurements tell us beyond the realized sample?
A corpus contains 600 relative-clause tokens from 30 speakers. A relative clause modifies a nominal expression, as in the book that Kim read.
The file proportion, the average speaker proportion, and a population proportion are different quantities.
One dataset is only one possible result of a sampling process.
First name the population quantity. Then define the estimation rule and study that rule under a stated sampling model.
Targets, observed units, and repeated measurements
Thirty Rochester speakers each produce twenty vowel tokens.
The table has 600 rows and 30 sampled speakers.
Buchstaller and Khattab ask how a sample of speakers can support a claim about a speech community.
Buchstaller and Khattab 2013
A target population contains the units about which we want to make a claim.
A sample contains the units observed through a particular selection process.
Let denote the population distribution. An estimand
is the population quantity selected by the linguistic question, such as a population mean or contrast.
A parameter, such as , is a fixed numerical feature of the modeled population.
An estimator is a random variable such as
Its realized value is an estimate of .
Six hundred tokens can sharpen summaries for 30 observed speakers. They do not create 600 independently sampled speakers.
Shared distributions and independent sampling units
Let when clause token contains an overtly expressed subject and otherwise. The realized values are
Random variables are independent and identically distributed (IID) when (i) they have the same distribution and (ii) their joint probability factors into marginal probabilities.
This claim concerns the sampling units, not merely distinct rows.
Recall the Bernoulli distribution, probability factorization, and independence.
For the observed sequence,
Distinct rows need not be independent. Clauses can share speakers, documents, genre, and discourse context.
An empirical proportion and its possible population target
The Universal Dependencies English Web Treebank teaching extract contains 12,280 personal-pronoun tokens. An older preprocessing rule labels case-distinct forms by surface form; for syncretic you and it, it uses the basic dependency relation as a fallback.
This heuristic accusative label is not the UD morphological feature Case=Acc.
UD English EWT
Heuristically accusative tokens:
Other tokens:
Define the descriptive file proportion
Without a population sampling model, describes the heuristic labels in this extract. It does not estimate the proportion of current UD Case=Acc annotations, explain case licensing, or justify independence across documents and repeated forms.
Fixed observations, varying parameter values
Which heuristic-label probability makes the observed token sequence most probable under the working model?
The likelihood function evaluates the observed values under each parameter value:
is a PMF over possible samples for fixed .
compares parameter values for fixed observations.
The log likelihood turns a product into a sum.
The maximum-likelihood estimator is the rule
For this sample, the realized maximum-likelihood estimate is .
Maximum likelihood selects a parameter value under the declared model. It does not establish that the model is adequate.
A positive dependency-distance response
The English Web Treebank teaching extract contains 2,594 selected wh-form tokens and their annotated basic-dependency heads. For token , let
Thus for adjacent tokens; zero is excluded because a dependency cannot connect a token to itself.
This response is not a filler-gap distance, and the stored name gap_position denotes the annotated head token rather than a silent gap.
Silveira et al. 2014
Under the working model ,
Recall the Poisson distribution.
Setting the derivative to zero gives
The distance data exclude zero, and the Poisson mean variance equality may fail.
Estimation and adequacy are separate questions.
An estimation rule across possible samples
In a lexical-decision task, a participant judges whether a presented letter string is a word. Let for a correct response and for an error.
For this illustration, participant-item encounters are independent sampling units and population accuracy is .
Each possible sample produces a possibly different sample-proportion estimator
The sampling distribution is the probability distribution of an estimator across possible samples.
The parameter is fixed. The estimator varies across samples. The estimate is the value for one realized sample.
Larger independent samples concentrate the sampling distribution around the same parameter.
A sampling distribution contains possible estimates, not possible population parameters.
Real lexical-decision experiments usually repeat participants and items. Then the independent-unit assumption must reflect both sources of clustering.
The spread of a sampling distribution
How much would the estimator’s value vary across repeated samples under the model?
The standard error is the standard deviation of an estimator’s sampling distribution.
For an IID sample mean,
Doubling does not halve the standard error. Multiplying by four does.
Replacing the true independent unit with the row count makes the standard error too small.
Average error across possible samples
The bias of an estimator is
An unbiased estimator may miss the parameter in one realized sample.
Realized error is not estimator bias.
Unbiasedness alone does not rank estimators. Their sampling variation may differ.
Combining bias and sampling variance
Mean squared error averages squared estimation error.
A slightly biased estimator can have lower MSE if its sampling variance is sufficiently smaller.
Bias and variance answer different questions. MSE evaluates their combined squared error.
One procedure, many repeated samples
Suppose a perception study compares mean prosodic-prominence ratings in conditions and .
The population contrast is the estimand. The difference between sample means, , is the estimator.
Prosodic prominence may depend on several cues; it is not another name for duration or .
The estimate locates the observed contrast. The standard error describes how much the estimator would vary across samples under the model.
A confidence interval combines them through a procedure with declared repeated-sampling behavior.
Suppose
Thus the central event is
Multiplying by the positive standard error and rearranging gives
After sampling, these endpoints are fixed.
Write for probability under the sampling model indexed by the fixed parameter. A 95% coverage probability satisfies
This simulation uses a known standard error and an exactly normal sampling distribution. It illustrates coverage; it does not verify the prominence model’s assumptions.
After observing the data, the parameter either lies in or it does not. The model does not license a probability for that fixed event.
The interval also does not contain 95% of participant responses. Response variation and estimator uncertainty are different quantities.
Nominal coverage depends on the standard error, reference distribution, dependence structure, and response model.
When the normal interval is unreliable
A Wald interval can extend outside zero and one and can have poor coverage for small samples or probabilities near a boundary.
For successes in trials,
Write for probability under .
The resulting interval is .
“Exact” refers to the finite sample binomial calculation. It does not establish independence, a shared probability, or representative sampling.
Approximating a sampling distribution by resampling
Twelve speakers supply median fundamental-frequency () measurements. measures the rate of vocal-fold vibration in hertz and serves as an acoustic correlate of perceived pitch.
The target is the population median.
The nonparametric bootstrap treats the empirical distribution of the observed independent units as a stand-in for the population, then resamples those units with replacement.
A percentile bootstrap confidence interval uses quantiles of the bootstrap estimates:
If tokens repeat within speakers, resample speakers and retain each selected speaker’s tokens together.
The resampling unit must match the independent unit in the population claim.
Notes: bootstrap confidence intervals · Praat: frequency and
Comparing a discrepancy with a reference distribution
A null hypothesis states a parameter claim. An alternative hypothesis states the competing parameter values.
The random statistic
varies across possible samples. The observed sample produces .
The null distribution is the reference distribution of the test statistic under the null model.
For a one-sided upper-tail test,
A two-sided test needs a declared extremeness rule. A p value is neither an effect size nor the probability that the null hypothesis is true.
Unknown response variance
Let be one predeclared acoustic contrast from speaker . The speakers, rather than the underlying token rows, are the independent sampling units.
The target is the population mean contrast .
The one-sample t-test applies the t procedure to these speaker contrasts.
For , ms, and ms,
The two-sided p value is .
Many tokens from one speaker do not provide many independent speaker contrasts.
Constructing a within-pair response
Every speaker produces one token in each of two vowel categories. A vowel category is a linguistic category; its acoustic measurements are numerical responses.
The measurements share vocal tract anatomy and recording conditions.
A paired design forms one within-pair difference for each independent pair.
Shared additive speaker shifts cancel inside .
Covariance determines the precision gained by pairing.
Positive within-pair covariance reduces the variance of the difference.
Sorting the two response vectors independently destroys the pairing and can create invalid differences.
Applying the one sample procedure to differences
Estimate the population mean of the speaker differences, .
The paired t-test is a one-sample t-test applied to .
Speaker differences use the Hillenbrand et al. first-formant () frequency estimates. is an estimated center frequency of the first vocal-tract resonance, measured in hertz; it is not a phoneme or vowel-category label.
Hillenbrand et al. 1995
For Hz,
The interval describes the mean within speaker vowel contrast, not the distribution of individual vowel measurements.
Different units in each group
How does uncertainty change when the two condition means come from different speakers?
Welch’s two-sample t-test allows the two groups to have different variances.
For and ,
The 95% interval is Hz.
Do not use an equal variance assumption merely because a software default or familiar formula includes it.
Paired and independent designs use different sources of variation. The test must match the design.
A conditional comparison for a two by two table
This teaching sample classifies noun lexemes by singular-lemma syllable count and by plural realization class: productive -(e)s versus another realization.
| productive -(e)s | other realization | total | |
|---|---|---|---|
| monosyllabic | 18 | 12 | 30 |
| polysyllabic | 26 | 4 | 30 |
| total | 44 | 16 | 60 |
The question is whether the two classifications are independent.
Fisher’s exact test constructs a conditional null distribution over tables with the observed margins fixed.
With the margins fixed, the upper-left count determines the other three cells. Under conditional independence,
The two-sided p value also includes other tables at least as improbable under the declared ordering rule.
The two-sided p value is approximately .
The observed table probability is only one term in that sum.
The odds of productive -(e)s are for monosyllabic lexemes and for polysyllabic lexemes. Their sample odds ratio is
Association between syllable count and plural realization class does not identify a causal or linguistic mechanism.
An approximate reference distribution
A corpus classifies 300 referring-expression tokens by grammatical role and expression form: personal pronoun or lexical noun phrase.
| personal pronoun | lexical noun phrase | total | |
|---|---|---|---|
| subject | 120 | 30 | 150 |
| object | 80 | 70 | 150 |
| total | 200 | 100 | 300 |
The chi-squared test of independence compares observed cell counts with counts expected under independence.
For cell ,
Thus the subject-pronoun cell has but .
For this table, with .
The total statistic records aggregate discrepancy across the whole table. Signed cell differences show which forms occur more or less often than the independence model expects.
Small expected counts weaken the approximation. Repeated speakers or documents also violate the independent row assumption.
Updating uncertainty about a parameter
A listener classifies an ambiguous plural-subject sentence. Under a collective reading, the predicate holds of the group; under a distributive reading, it holds of each relevant member.
Let for a collective response and for a distributive response.
Experimental evidence on collective and distributive readings
A prior distribution assigns probability to parameter values before the present data are observed.
Here denotes the response PMF and denotes the parameter density:
Recall Bayes’ rule.
The posterior distribution represents parameter uncertainty after conditioning on the observed data.
The posterior is a probability distribution over the parameter. The likelihood alone is not.
Claims from a posterior distribution
Suppose the analysis produces draws
Their empirical distribution approximates only when the computational procedure is valid.
A posterior summary is a numerical description calculated from the posterior distribution or its draws.
State whether the claim concerns direction, magnitude, or a predicted response; then select the corresponding summary.
For draws ,
The positive tail pulls the mean upward.
A loss function assigns a cost to estimation error.
Neither is automatically the Bayesian estimate.
If the and posterior quantiles are
then of the fitted posterior probability for lies between them, conditional on the likelihood, prior, and observed data.
Finite-draw endpoints have Monte Carlo error. And posterior probability does not establish that the model or prior represents the linguistic sampling process.
Numerically similar confidence and credible intervals can thus support different interpretations.
Let map each outcome to a response-probability difference. For a directional claim, define
Let equal 1 on and 0 otherwise. Then estimate
This value is neither a p value nor a measure of association size.
For odds, compute before taking summaries. In general,
For each paired draw, compute
Separate marginal intervals discard which values occurred together and do not determine the contrast distribution.
A parameter summary, predicted-response contrast, and directional probability answer different questions. Predeclare diagnostic profiles or show a complete small grid.
Which probability statement does support, and why must a multi-parameter contrast be computed within each draw?
An analytic Beta-binomial update
A phonological study codes whether word-final /t/ or /d/ is absent in a predeclared environment. It observes 14 deletion labels and 6 realized-stop labels.
The shared-rate model below suppresses known conditioning by segmental context, morphological class, and speaker.
A conjugate prior yields a posterior in the same distribution family after multiplication by the likelihood.
The update resembles adding prior successes and failures. This pseudocount analogy explains the algebra; the shape parameters are not literal unobserved tokens.
Notes: conjugate priors · Module 2: Beta distributions · Module 2: binomial distributions
Checking implications before observing the present data
What response counts does the complete prior model expect for ten ambiguous sentences?
The prior predictive distribution averages the observation model over the prior.
Draw one parameter value for a complete simulated dataset, then generate the responses conditional on that value.
Inspect predictions on the response scale. Ask whether the implied counts are plausible before using the present responses.
Drawing a new response probability for every trial represents a different sampling structure from drawing one for the complete ten trial dataset.
Predicting after the update
The posterior predictive distribution averages the observation model over posterior uncertainty.
Draw from the posterior, then draw the future value from the observation model conditional on .
Predictive spread contains parameter uncertainty and future response variation.
These are distinct sources of variation.
Prediction asks what future observations may look like. Model checking compares replicated observations with the observed data.
The computational problem
A stop’s closure duration is the time interval during which oral airflow is blocked. Four independent speakers supply standardized closure-duration contrasts; each contrast is dimensionless after standardization.
The normalizing constant is
Define the posterior kernel
Ratios of do not require .
Computational difficulty does not imply broad posterior uncertainty. A concentrated posterior can still have an intractable normalizer.
How can we calculate posterior expectations without evaluating every possible parameter value?
Replacing an expectation with a sample average
Let and suppose
The target is .
Define by
If when and otherwise, then
Draw independently from the target:
For a continuous density and integrable ,
For a discrete target, replace the integral with a sum.
Monte Carlo error is variation caused by using a finite number of simulation draws.
For an indicator estimator,
Multiplying by four tends to divide MCSE by two.
More simulation draws reduce Monte Carlo error. They do not reduce posterior uncertainty.
More draws are not more linguistic evidence.
Which uncertainty can be reduced without collecting more linguistic observations?
Sampling from a different distribution
Direct draws from the target distribution may be unavailable.
Draw from a proposal distribution with density and reweight the draws toward the target density .
The importance weight is
when the target is normalized and the proposal covers its support.
A few very large weights can dominate the estimate. Unweighted draw frequency comes from ; weighted averages approximate the target with density .
Dependent simulation draws
A Markov chain is a sequence of random states. Given the current state, the distribution of the next state does not depend on the earlier path.
For a permitted (measurable) set of possible next-state values,
A distribution is stationary when one transition leaves it unchanged.
Irreducibility means that the chain can eventually reach every region that the target distribution assigns positive probability.
Aperiodicity rules out deterministic cycling through states at a fixed period.
These conditions concern long-run access. They do not guarantee efficient exploration in a finite run.
Using a chain to approximate a target distribution
For a standardized articulation-rate contrast ,
The likelihood and prior define this target.
Markov chain Monte Carlo (MCMC) constructs a Markov chain whose stationary distribution is the target.
are intended to have long-run density .
Under conditions that make the chain ergodic,
Dependence does not invalidate the average. The transition must preserve the target, and the chain must reach its relevant regions.
Consider
Its stationary distribution is ; initialize at .
The first 1,000 states are omitted because initialization was deliberately extreme. No observable iteration makes a finite chain exactly stationary.
Multiple chains from dispersed plausible values should move toward and repeatedly explore the same region.
Burn-in omits an initialization period under a fixed transition rule.
Warmup in an adaptive sampler also tunes quantities such as step size and mass matrix; its changing transition states are not posterior draws.
| object | role |
|---|---|
| posterior target | distribution the analysis seeks to approximate |
| transition distribution | conditional rule for moving to a new state |
| empirical draw distribution | finite states produced by one run |
A wide posterior may reflect limited linguistic information. A slowly moving chain reflects computational inefficiency.
More iterations can reduce Monte Carlo error without adding linguistic evidence or narrowing the posterior.
Propose, compare, and accept
The Metropolis-Hastings algorithm corrects proposal moves with an acceptance probability.
For a symmetric proposal and standard normal target, a move from to has
The shared normalizing constant cancels.
A proposal that is too narrow moves slowly. A proposal that is too wide is rejected often.
Using gradients to propose distant states
A local proposal may bounce slowly through a narrow, correlated high-density region.
The gradient supplies local slope. An auxiliary momentum variable supplies direction, allowing a proposal to travel before the accept-or-reject step.
Hamiltonian Monte Carlo (HMC) augments the parameter with momentum and follows a numerically approximated Hamiltonian trajectory.
Potential energy is the negative log target kernel. Kinetic energy uses auxiliary momentum and a positive-definite mass matrix :
The Hamiltonian is the sum of potential and kinetic energy.
The gradient updates momentum using the local log-density slope.
For step size ,
For , , and ,
A Metropolis correction accounts for numerical error.
HMC uses local gradient information to make long proposals with high acceptance probability.
Matching movement to scale and correlation
The mass matrix determines the kinetic energy and maps momentum into parameter velocity.
A suitable adjusts movement to posterior scales and correlations.
The mass matrix is symmetric positive definite. A unit mass matrix uses one scale for every direction. A diagonal mass matrix adapts scales. A dense mass matrix also adapts correlations.
Stan estimates the mass matrix from warmup draws as an inverse covariance approximation. The learned geometry changes computation, not the model’s posterior target.
Choosing a trajectory length
The No-U-Turn Sampler (NUTS) builds an HMC trajectory and stops extending it when the path begins to reverse.
Grow a balanced trajectory tree, check the turning condition, and select one valid state from the path.
NUTS chooses trajectory length. Warmup separately adapts step size and the mass matrix.
Stopping at the first rejected leapfrog point would bias the trajectory. NUTS selects from a valid constructed path.
Notes: the No-U-Turn Sampler · Stan Reference Manual: HMC and NUTS
Translating a factorization into a program
Stan is a probabilistic programming language that uses gradient-based MCMC for continuous parameter models.
The model block should match the declared joint factorization.
Parameter constraints imply transformations and Jacobian adjustments handled by Stan.
A program can compile and sample from the model that was coded while the model remains inappropriate for the linguistic data.
Visual evidence about chain behavior
A trace plot displays saved parameter draws against iteration for each chain.
Look for stable location, overlap across chains, repeated crossing, and continued movement.
Drift indicates nonstationarity. Separated chains indicate incomplete exploration or distinct regions.
Overlapping traces do not prove convergence. They provide one diagnostic view that must agree with numerical summaries.
Comparing within and between chain variation
If chains explore the same target, variation within each chain should resemble variation between chains.
The statistic, specifically the rank-normalized split , formalizes this comparison after splitting chains and rank-normalizing their draws.
near one indicates that the chains have similar location and scale for the monitored parameter.
Chains can agree inside the same incomplete region. is not a certificate that the posterior was fully explored.
Inspect the largest across all reported and scientifically relevant parameters, not one convenient coefficient.
Notes: cross-chain agreement · Stan: convergence diagnostics
Dependence across saved iterations
Autocorrelation is correlation between draws separated by a fixed lag within one stationary chain.
Let denote correlation between random variables and under the stationary distribution. Then
For , .
Autocorrelation can reduce Monte Carlo efficiency even when the chain has the correct stationary distribution.
Thinning discards draws and usually does not improve estimation for a fixed computational run.
Information in dependent draws
Effective sample size is the number of independent draws that would provide comparable Monte Carlo precision.
For a scalar posterior quantity , define its lag- autocorrelation as . Then, for one stationary chain,
Positive autocorrelation usually lowers ESS; negative autocorrelation can raise it above .
If and , then
Bulk ESS and tail ESS are summary-specific refinements. Bulk ESS concerns central summaries; tail ESS concerns quantiles and tail probabilities.
Effective sample size counts information in simulation draws. It does not count speakers, items, or experimental observations.
Numerical failure in difficult geometry
A divergent transition occurs when numerical integration error becomes too large along an HMC trajectory.
A positive scale parameter and latent deviations can create a narrow, highly curved posterior region.
This shape is called a funnel.
Locate divergent draws in parameter space. Then reconsider parameterization, scale, prior information, and model geometry.
Discarding divergent draws changes the empirical distribution and does not repair the sampler’s failure to explore the target.
Locating difficult posterior geometry
A pairs plot displays pairwise projections of posterior draws.
Inspect funnels, narrow ridges, strong correlations, multimodal separation, and the locations of divergent draws.
A geometric association between parameters does not show that one parameter causes the other.
Pairs plots connect a sampler warning with the posterior region where computation is difficult.
Calculating within posterior draws
The generated quantities block computes derived values or simulations after each saved parameter draw.
Compute a deterministic contrast within each draw to preserve joint uncertainty.
Simulate a replicated response when the quantity is stochastic.
The block does not change the posterior target because its calculations occur after parameter sampling.
Preserve participant, item, and observation indices when simulated quantities must retain the fitted dependence structure.
Notes: generated quantities · Stan Reference Manual: program blocks
Comparing observations with replicated data
A posterior predictive check compares observed data with replicated data generated under posterior draws.
For each posterior draw, generate the replicated dataset and retain its realization using the observed design and indexing.
Recall the posterior predictive distribution.
A discrepancy function reduces observed and replicated data to a feature the model should reproduce.
Observed low-frequency decision-time standard deviation: ms
Replicated 95% interval: ms
Only replications reach ms.
Compare observed and replicated variance, the proportion at a response-scale endpoint, group means, or the frequency of very long word types.
Checking only a statistic that the fitted model must reproduce closely can miss failures elsewhere in the distribution.
A mismatch identifies a feature the fitted model fails to reproduce. Agreement does not establish that the model is true.
Notes: posterior predictive checks · Stan User’s Guide: predictive checks
From population target to model check
What can observed speakers, tokens, items, or measurements tell us beyond the realized sample?
Every result must identify its population target, sampling model, procedure, computational check, and fitted-model check.
Targets and procedures · Posterior distributions · Posterior predictive checks