Populations and samples

The foundations chapter asked what one observation represents. We now ask the next question: which population are those observed units meant to represent? Suppose we record 30 speakers in Rochester and measure the duration of one local vowel category. Duration is an acoustic measurement on each vowel token, while speaker is the unit recruited into the study.

Possible answers include current residents of Rochester, lifelong residents, speakers in a particular age range, or only the 30 people who participated. The spreadsheet cannot choose among these answers. The research question and recruitment process must do that work.

A target population is the collection of units about which an inference is intended, while a sample is the collection observed through a particular selection process. The target is part of the research question, not every unit that happens to resemble the observed ones.

NoteReading: Buchstaller and Khattab (2013), Chapter 5

Buchstaller and Khattab, Population samples, in Research Methods in Linguistics.

Why this reading? The chapter asks how sampled speakers can support a claim about a speech community. It compares sampling strategies and shows why a large number of tokens cannot repair a mismatch between the recruited speakers and the population stated in the conclusion.

Specifying the population unit

A population must be a population of something, so before we count rows, we must name that unit. It might contain

  1. speakers in a speech community;
  2. utterances produced in a genre;
  3. lexical types represented by a dictionary;
  4. sentence items instantiating a construction;
  5. judgments supplied by a class of comprehenders.

These units are not interchangeable. Ten thousand tokens from five speakers provide many utterance observations but only five sampled speakers.

If the target concerns speakers, the speaker count supplies the direct replication at that level. If the target concerns tokens from those five speakers during the recorded sessions, the larger token count may be relevant. The unit specified in the population claim determines which count matters.

Describing the route into the sample

Now ask how the observed units entered the dataset. Speakers might be randomly selected from a known roster, recruited through community organizations, drawn from an available student pool, or self-selected after an advertisement.

Each route can support some descriptions. None automatically supports every claim about “English speakers.”

For the Rochester study, a precise statement might be:

The target population is adult speakers who grew up in the Rochester metropolitan area and currently live there. The sample contains 30 volunteers recruited through neighborhood organizations during Fall 2026.

This statement makes two objects inspectable: which speakers the conclusion concerns and which speakers had a route into the sample.

Defining the population target

Suppose the target is mean vowel duration among speakers in the population. Write the unknown population mean as

\mu.

A fixed numerical feature of the modeled population is a parameter. Here the parameter \mu is also the estimand introduced in the foundations chapter. A model may contain other parameters that are needed for the analysis but are not themselves the estimand.

From one observed sample of n speakers, we calculate

\bar{x}=\frac{1}{n}\sum_{i=1}^{n}x_i.

This realized sample mean is an estimate. It is a number computed from the observed sample and used to learn about \mu.

The roles differ:

object role
target population declares which units the claim concerns
population parameter \mu fixed target under the frequentist model
realized sample units that entered through the selection process
estimate \bar{x} numerical summary of this sample

Later pages introduce the estimator as a rule across all possible samples. For now, the main distinction is between the population target and the observed sample summary.

Working through two candidate targets

Imagine each of the 30 speakers produces 20 vowel tokens. The dataset has 600 token rows.

One target could be the mean duration of tokens produced by these 30 speakers in this task. Another could be the mean of speaker-level average durations among adult Rochester speakers represented by the recruitment process. These are different estimands because the first weights tokens and the second weights speakers.

For the second target, one useful estimate first averages tokens within speaker and then averages the 30 speaker summaries. The 600 tokens improve each observed speaker summary, but they do not create 600 independent speakers.

Rows may multiply at one level while the target population varies at another. The relevant amount of replication thus depends on the population target.

Inspecting the levels in base R

Code
speaker <- rep(sprintf("s%02d", 1:30), each = 20)
token <- sequence(rep(20, 30))

length(token)
length(unique(speaker))

The two results are 600 token observations and 30 sampled speakers. Neither number replaces the other.

A large sample can target the wrong population

More observations may reduce sampling variability under a model. They do not broaden recruitment. A study of 5,000 university students may describe those students precisely while providing little direct information about speakers over age 70.

A large row count cannot substitute for a justified connection to the target population. Population definition must precede estimator evaluation because precision about the wrong target does not repair the target.

Check your understanding

  1. A corpus contains 50,000 tokens from four speakers. Give one target for which 50,000 is the relevant observation count and another for which four is the relevant sampled-unit count.
  2. Distinguish the population parameter \mu from the realized sample mean \bar{x}.
  3. What information about recruitment is needed before generalizing from volunteers to a speech community?
  4. Why can a very large convenience sample remain poorly connected to the intended population?

Once the population and sample units are specified, we can state a simple model for repeated observations. The next page introduces independent and identically distributed observations.