Paired observations
The one-sample procedure requires one response from each independent unit. A paired design creates that response by taking a within-pair difference. Consider an acoustic study in which every speaker produces two vowel categories. The response is first-formant frequency, written F_1 and measured in hertz. F_1 is a resonance of the vocal tract that is commonly interpreted in relation to vowel height; it is not pitch or fundamental frequency. The two tokens are distinct utterances, but they share a speaker’s vocal tract anatomy, recording environment, and many features of speaking style.
A paired design links two measurements because they arise from the same unit or from units matched before analysis. The pair, rather than either measurement alone, carries the contrast of interest, raising the question of how to turn those two linked measurements into one response at the independent-unit level. We do so by taking a within-pair difference.
From two measurements to one contrast
Let Y_{sA} and Y_{sB} be measurements from conditions A and B for pair s, where s=1,\ldots,S. Define
D_s = Y_{sA}-Y_{sB}.
The sign is a convention, but it must be fixed before interpretation. If A is a high vowel and the response is F_1, a negative D_s means that the high vowel has lower F_1 than the comparison vowel for speaker s.
After differencing, the dataset contains S contrast observations:
| pair | Y_A | Y_B | D=Y_A-Y_B |
|---|---|---|---|
| 1 | 312 | 381 | -69 |
| 2 | 347 | 405 | -58 |
| 3 | 290 | 372 | -82 |
| 4 | 401 | 438 | -37 |
The original measurement unit may be a trial or token, but the unit for the condition contrast is now the pair. Four pairs provide four observed differences. Counting the table as eight unrelated values would double the apparent number of contrast-bearing units.
What differencing removes
Suppose each measurement can be decomposed as
Y_{sc}=\alpha_s+\delta_c+\varepsilon_{sc},
where \alpha_s is a stable pair-specific baseline, \delta_c is the condition contribution, and \varepsilon_{sc} is remaining measurement variation. The difference is
D_s = (\alpha_s+\delta_A+\varepsilon_{sA}) - (\alpha_s+\delta_B+\varepsilon_{sB}) = \delta_A-\delta_B+\varepsilon_{sA}-\varepsilon_{sB}.
The baseline \alpha_s cancels. In the acoustic example, speakers may differ substantially in overall F_1 because their vocal tracts differ. Those baseline differences are not evidence about the within-speaker vowel contrast, so removing them can sharpen the comparison.
The same logic applies across linguistic designs:
- the same participant judges two versions of an item;
- the same lexical item appears in two syntactic frames;
- the same language is measured before and after a coding revision; or
- two corpus samples are matched on speaker and lexical context.
The paired unit must follow from the design or an explicit matching rule. Similar-looking rows are not automatically pairs.
Covariance determines the change in precision
Now, you might wonder whether differencing always makes an estimate more precise. The variance and covariance rules from the preceding module give
\operatorname{Var}(Y_A-Y_B) = \operatorname{Var}(Y_A) + \operatorname{Var}(Y_B) - 2\operatorname{Cov}(Y_A,Y_B).
When the two measurements are positively associated within pairs, the covariance term reduces the variance of the difference. Stable speaker baselines, for instance, tend to make both vowel measurements high or both low. The paired contrast subtracts that shared displacement.
Pairing does not guarantee a smaller variance. Negative within-pair covariance would increase the variance of D_s, and uninformative matching may produce covariance near zero. The design still determines the appropriate unit. Precision is a consequence of that design, not the criterion used to decide whether pairing exists.
Whether observations are paired is determined by the design, not by which analysis produces a narrower interval. The observations were paired before their values were known because the same speaker, item, or matched unit generated them.
Pairing versus repeated measurements
A simple pair contains two measurements. Linguistic datasets often contain more: twelve vowels from each speaker, forty sentences from each participant, or several annotations of each predicate and argument pair. A two-condition question can sometimes be reduced to one difference per repeated unit, as above. But a dataset with several conditions, missing cells, or multiple tokens in each condition requires additional decisions.
For instance, if each speaker produces five tokens of A and five of B, one could average within each speaker and condition before differencing. That produces one difference per speaker, but discards token variation. A later multilevel model can retain the tokens while representing the speaker structure. Thus, the simple paired difference is a complete analysis for one measurement per condition, and a useful design diagnostic for more complex repeated data.
Incomplete and invalid pairs
If one measurement is missing, its pair does not yield a difference, so complete-pair analysis targets units observed in both conditions. This target may differ from the target over all enrolled units when missingness is related to condition or response. A report should give the number of original units, complete pairs, and exclusions.
Arbitrary post hoc matching is a different problem. Pairing each short word with a long word that happens to have a similar frequency may create a useful matched design, but the matching variables and algorithm must be fixed and reported. Choosing matches that maximize a desired result makes the pairing outcome-dependent.
Pair observations because the same unit generated both measurements or because a prespecified matching design linked the units. Do not infer pairing from a convenient row order or from the interval produced after fitting.
Checking the inferential claim
- An experiment records one judgment for each of two sentence versions from every participant. What is the pair identifier, and what does one difference represent?
- Each participant sees version A of one item and version B of a different item. Are the two trials paired by participant for a claim about the item contrast? What additional structure should be considered?
- If \operatorname{Cov}(Y_A,Y_B)>0, why is the variance of the within-pair difference smaller than the sum of the marginal variances?
What pairing accomplishes
A paired design makes the within-pair difference the observation that carries a two-condition contrast. Differencing removes a stable pair baseline, and positive within-pair covariance can increase precision. The next page uses the distribution of these differences to estimate an average contrast and quantify its uncertainty.