The paired design page used covariance between linked measurements. When two groups contain different independently sampled speakers, that covariance is absent. Suppose one group of speakers produces vowels in dialect A and a different group produces vowels in dialect B. The response is first-formant frequency, written F_1 and measured in hertz. It is a vocal-tract resonance commonly interpreted in relation to vowel height, not a measure of pitch.
\Delta=\mu_A-\mu_B.
The estimator is the difference between the two sample means:
\widehat{\Delta}=\overline{Y}_A-\overline{Y}_B.
Because the groups contain different speakers, there is no within-speaker difference to calculate. The question is how uncertainty from the two mean estimators combines, and under independent sampling, both contributions remain.
Adding independent sampling variances
For independent groups, the general variance rule contains a covariance term, but independence makes that term zero. Thus
This is the standard error in Welch’s two-sample t-test, which allows the two variance contributions to differ. The test does not assume that the population variances are equal.
The value need not be an integer. For these data, \nu\approx23.005. The two-sided p value is approximately .150, and the 95% confidence interval is approximately
[-95.51,15.51]\text{ Hz}.
The interval includes contrasts in both directions. These data do not estimate the direction precisely under this procedure.
For row-level data with variables f1 and dialect, R uses Welch’s two-sample t-test by default:
t.test(f1 ~ dialect, data = vowel_tokens)
Unequal variances
Student’s two-sample t-test assumes \sigma_A^2=\sigma_B^2 and pools the sample variances before calculating a standard error. That equality is an additional model claim. It does not follow from having two groups.
Welch’s two-sample t-test remains appropriate when the variances differ and generally sacrifices little precision when they happen to be equal. A preliminary variance test does not supply a linguistic reason to adopt the pooled model.
Matching the procedure to the design
If the same speakers produced both dialect styles, the observations would be paired, and treating them as independent would discard within-speaker covariance. Conversely, matching unrelated speakers by row number would invent pairing.
The layout of a spreadsheet does not determine whether observations are paired. Recruitment and measurement determine whether observations are paired or independent.
NoteProblem-set reading
PS1, Task 4: diagnose the independence error compares an independent-samples analysis with the paired analysis required by the recording design. It is included because the same measurements make the difference between those procedures concrete.
Check your understanding
Why are the two sampling variance contributions added?
Which group contributes more uncertainty in the worked example?
What extra assumption does the pooled procedure make?
How would the data structure change if every speaker produced both conditions?
---title: "Comparing two independent means"---The [paired design page](paired-observations.qmd) used covariance between linked measurements. When two groups contain different independently sampled speakers, that covariance is absent. Suppose one group of speakers produces vowels in dialect A and a different group produces vowels in dialect B. The response is [first-formant frequency](https://www.fon.hum.uva.nl/praat/manual/Sound__To_Formant__burg____.html), written $F_1$ and measured in hertz. It is a vocal-tract resonance commonly interpreted in relation to vowel height, not a measure of pitch.$$\Delta=\mu_A-\mu_B.$$The estimator is the difference between the two sample means:$$\widehat{\Delta}=\overline{Y}_A-\overline{Y}_B.$$Because the groups contain different speakers, there is no within-speaker difference to calculate. The question is how uncertainty from the two mean estimators combines, and under independent sampling, both contributions remain.## Adding independent sampling variancesFor independent groups, the general variance rule contains a covariance term, but independence makes that term zero. Thus$$\begin{aligned}\operatorname{Var}(\widehat{\Delta})&=\operatorname{Var}(\overline{Y}_A-\overline{Y}_B)\\&=\operatorname{Var}(\overline{Y}_A)+\operatorname{Var}(\overline{Y}_B).\end{aligned}$$The minus sign disappears because multiplying a random variable by $-1$ does not change its variance. Independence makes the covariance term zero.Estimate each population variance from its group:$$\widehat{\operatorname{SE}}(\widehat{\Delta})=\sqrt{\frac{s_A^2}{n_A}+\frac{s_B^2}{n_B}}.$$This is the standard error in [**Welch's two-sample t-test**](https://stat.ethz.ch/R-manual/R-devel/library/stats/html/t.test.html), which allows the two variance contributions to differ. The test does not assume that the population variances are equal.## Calculating an $F_1$ contrastSuppose the group summaries are$$\overline{y}_A=510\text{ Hz},\quads_A=60\text{ Hz},\quadn_A=20$$and$$\overline{y}_B=550\text{ Hz},\quads_B=90\text{ Hz},\quadn_B=15.$$The observed estimate is$$\widehat{\delta}=510-550=-40\text{ Hz}.$$The two estimated variance contributions are$$\frac{60^2}{20}=180$$and$$\frac{90^2}{15}=540.$$Thus$$\widehat{\operatorname{SE}}(\widehat{\Delta})=\sqrt{180+540}=\sqrt{720}\approx26.833\text{ Hz}.$$For $H_0:\Delta=0$, the observed statistic is$$t_{\mathrm{obs}}=\frac{-40-0}{26.833}\approx-1.491.$$## Approximating the degrees of freedomWelch's two-sample t-test uses$$\nu\approx\frac{\left(s_A^2/n_A+s_B^2/n_B\right)^2}{\frac{(s_A^2/n_A)^2}{n_A-1}+\frac{(s_B^2/n_B)^2}{n_B-1}}.$$The value need not be an integer. For these data, $\nu\approx23.005$. The two-sided p value is approximately $.150$, and the 95% confidence interval is approximately$$[-95.51,15.51]\text{ Hz}.$$The interval includes contrasts in both directions. These data do not estimate the direction precisely under this procedure.```{r}#| label: verify-welch-summary#| echo: truemean_a <-510mean_b <-550sd_a <-60sd_b <-90n_a <-20n_b <-15estimate <- mean_a - mean_bvariance_a <- sd_a^2/ n_avariance_b <- sd_b^2/ n_bse <-sqrt(variance_a + variance_b)df <- (variance_a + variance_b)^2/ (variance_a^2/ (n_a -1) + variance_b^2/ (n_b -1))t_observed <- estimate / sep_value <-2*pt(-abs(t_observed), df = df)interval <- estimate +c(-1, 1) *qt(.975, df = df) * sestopifnot(abs(se -26.832816) < .000001)stopifnot(abs(df -23.005405) < .000001)c(estimate = estimate, se = se, df = df, p = p_value,lower = interval[1], upper = interval[2])```For row-level data with variables `f1` and `dialect`, R uses Welch's two-sample t-test by default:```rt.test(f1 ~ dialect, data = vowel_tokens)```## Unequal variancesStudent's two-sample t-test assumes $\sigma_A^2=\sigma_B^2$ and pools the sample variances before calculating a standard error. That equality is an additional model claim. It does not follow from having two groups.Welch's two-sample t-test remains appropriate when the variances differ and generally sacrifices little precision when they happen to be equal. A preliminary variance test does not supply a linguistic reason to adopt the pooled model.## Matching the procedure to the designIf the same speakers produced both dialect styles, the observations would be paired, and treating them as independent would discard within-speaker covariance. Conversely, matching unrelated speakers by row number would invent pairing.The layout of a spreadsheet does not determine whether observations are paired. Recruitment and measurement determine whether observations are paired or independent.::: {.callout-note title="Problem-set reading"}[PS1, *Task 4: diagnose the independence error*](../problem-sets/ps1/ps1.qmd#task-4-diagnose-the-independence-error-20-points) compares an independent-samples analysis with the paired analysis required by the recording design. It is included because the same measurements make the difference between those procedures concrete.:::## Check your understanding1. Why are the two sampling variance contributions added?2. Which group contributes more uncertainty in the worked example?3. What extra assumption does the pooled procedure make?4. How would the data structure change if every speaker produced both conditions?