The preceding page defined the sampling distribution of an estimator, and we now ask how spread out that distribution is. Its standard error is the standard deviation of that sampling distribution:

\operatorname{SE}(\widehat{\Theta}) =\sqrt{\operatorname{Var}(\widehat{\Theta})}.

This definition concerns an estimator, not an individual response. We must keep those two objects separate.

Beginning with response variation

Suppose vowel duration in a population has mean \mu and standard deviation

\sigma=30\text{ ms}.

The standard deviation is the root mean squared distance of individual vowel durations from \mu. If we draw another vowel, its duration may differ substantially from the first one. This is variation among responses.

Now suppose that we repeatedly draw independent samples of 25 vowels and calculate the sample mean in every sample. Those sample means also vary, but they vary less than the individual durations. This is variation among estimates across samples.

quantity object that varies question answered
response standard deviation individual vowel duration X_i How dispersed are durations?
standard error sample mean \overline{X} How dispersed are mean estimates?

Both quantities have units of milliseconds. Their shared units do not make them interchangeable.

Deriving the standard error of a mean

The estimator of the population mean is

\overline{X}=\frac{1}{n}\sum_{i=1}^nX_i.

We first move the constant 1/n outside the variance. The square appears because multiplying a random variable by a constant multiplies its variance by the constant squared:

\operatorname{Var}(\overline{X}) =\frac{1}{n^2}\operatorname{Var}\left(\sum_{i=1}^nX_i\right).

If the observations are independent, their variances add. If they also have a common variance \sigma^2, then

\operatorname{Var}\left(\sum_{i=1}^nX_i\right) =\sum_{i=1}^n\operatorname{Var}(X_i) =n\sigma^2.

Substitution gives

\operatorname{Var}(\overline{X}) =\frac{n\sigma^2}{n^2} =\frac{\sigma^2}{n}.

Taking the square root yields the standard error:

\operatorname{SE}(\overline{X}) =\frac{\sigma}{\sqrt{n}}.

For 25 vowels with \sigma=30 ms, the calculation is

\operatorname{SE}(\overline{X}) =\frac{30}{\sqrt{25}} =6\text{ ms}.

The individual durations have a standard deviation of 30 ms, while means based on 25 independent durations have a standard error of 6 ms.

Changing the sample size

The square root in the denominator determines how precision changes with sample size.

n calculation standard error
25 30/\sqrt{25} 6 ms
100 30/\sqrt{100} 3 ms
400 30/\sqrt{400} 1.5 ms

Multiplying the sample size by four divides the standard error by two. Merely doubling the sample size does not halve it.

Code
sigma <- 30
n <- c(25, 100, 400)
se <- sigma / sqrt(n)

stopifnot(isTRUE(all.equal(se, c(6, 3, 1.5))))
data.frame(n, se)

Estimating an unknown standard error

Because the population standard deviation \sigma is usually unknown, let S be the sample-standard-deviation estimator. The corresponding estimated-standard-error rule is

\widehat{\operatorname{SE}}(\overline{X}) =\frac{S}{\sqrt{n}}.

After a sample is observed, S produces a realized sample standard deviation s. Suppose 25 observed vowels have s=32 ms. The realized estimated standard error is

\frac{32}{\sqrt{25}}=6.4\text{ ms}.

This number is an estimate of estimator spread rather than a revised estimate of response spread. The observed response standard deviation remains 32 ms.

Recognizing the independence failure

The derivation used independence when it added n separate variances. Repeated tokens from one speaker may be correlated. Counting 400 such tokens as 400 independent observations can thus produce an estimated standard error that is much too small.

The calculation treats the number of rows as the amount of independent information. The relevant n is instead determined by the sampling design and the dependence structure. A large file can contain little independent information.

Check your understanding

  1. If \sigma=40 ms and n=16, what is \operatorname{SE}(\overline{X})?
  2. What sample size gives a standard error of 5 ms when \sigma=40 ms?
  3. Why can response standard deviation and standard error have the same units but different interpretations?
  4. Which step of the derivation fails when observations are positively correlated?