Learning from a sample

Suppose a corpus contains 600 relative clauses produced by 30 speakers. A canonical headed relative clause is a clausal modifier related to a constituent, typically a noun phrase; in a subject relative, that constituent is interpreted as the subject of the relative clause. We can calculate the proportion of subject relatives among the 600 tokens. We can also calculate a proportion for each speaker and then average the 30 speaker proportions. These two summaries need not agree because the first gives equal weight to tokens and the second gives equal weight to speakers.

Neither summary is automatically a population quantity. If the research question concerns speakers in a region, the population includes speakers who did not contribute to the corpus. The question for this module is thus not merely how to calculate a sample summary. It is: what can the observed speakers and clauses tell us about a specified quantity in a specified population?

Populations, samples, and estimators

We begin by distinguishing populations from samples. The population must identify the units to which the claim applies, while the sample contains the units that were observed. An independent sampling model then states how repeated observations are related to a population parameter.

The pronoun case example uses a binomial model for one corpus proportion. Maximum likelihood selects the parameter value that assigns the greatest probability to the observed sample. An estimator is the rule that maps each possible sample to an estimate, while the estimate is the value obtained from the realized sample.

The sampling distribution of an estimator allows us to define its standard error, bias, and mean squared error. Confidence intervals and exact binomial intervals are procedures whose coverage is defined across repeated samples. Null hypotheses and test statistics compare an observed statistic with a reference distribution under a stated null hypothesis.

Later pages apply these definitions in increasing order of design complexity: first to one mean, then to paired acoustic measurements, then to two independent means, and finally to contingency tables. In each case, we will name the independent unit before calculating an interval or test.

Updating uncertainty with Bayes’ rule

Bayesian inference begins with a probability model for the observations and adds a probability distribution over unknown parameters. The posterior distribution page uses Bayes’ rule to combine a prior distribution and likelihood. The posterior summaries page then asks which features of that distribution address a particular linguistic claim.

Some posterior distributions can be calculated algebraically. Conjugate priors make one class of such calculations possible by keeping the posterior in a known family.

The next distinction concerns prediction. A prior predictive distribution states what the model predicts before the present observations update it, while a posterior predictive distribution states what it predicts after that update. Both distributions put the prediction on an observable response scale. Other models leave a posterior that is mathematically defined but not available as a familiar normalized formula. We thus begin with a finite update, summarize it, derive a conjugate update, predict before and after conditioning, and only then turn to numerical approximation. This order keeps the inferential target separate from the algorithm used to approximate it.

Posterior computation and checking

The computational sequence also has an order. Monte Carlo integration first explains why averages of target draws approximate expectations. Importance sampling changes the distribution used to produce those draws. Markov chains and Markov chain Monte Carlo then permit dependent draws when direct independent sampling is unavailable.

The remaining pages move from Metropolis-Hastings transitions to Hamiltonian Monte Carlo and Stan’s implementation. Only after the sampler has been defined do we read trace plots, \widehat R, autocorrelation, effective sample size, divergences, and parameter pairs. Generated quantities and posterior predictive checks then return us to the response scale. This final return matters: a sampler may accurately approximate a posterior for a model that predicts the linguistic data poorly.

For each inferential result, identify the population quantity, the sampling or observation model, and the procedure that maps the realized sample to the reported value. In the relative clause example, a token proportion, an average speaker proportion, and a regional population proportion are three different quantities even when they are calculated from the same file.

The chapter dependency graph shows the probability concepts used in this module. We assume the earlier definitions of conditional probability, independence, expectation, and variance. Begin with populations and samples.