Confidence intervals

The preceding pages separated a point estimate from its standard error. A confidence interval is a procedure that combines those two quantities to produce a range of parameter values.

The word procedure matters. Confidence is a long-run property of the rule that generates intervals, not a probability attached to the fixed endpoints observed in one study. We will first derive the rule while the data are still random and only then interpret the interval produced by the observed sample.

Identifying the target and estimate

Suppose a perception study compares mean prosodic-prominence ratings in two conditions. Prominence concerns the extent to which a word or syllable stands out relative to surrounding material. It is a perceptual and phonological notion that may depend on several cues, not another name for duration or fundamental frequency. The population contrast is

\Delta=\mu_A-\mu_B.

This parameter is the estimand, while the difference between sample means, \widehat{\Delta}, is the estimator. The observed data produce the estimate

\widehat{\delta}=.42

rating points. Its estimated standard error is

\widehat{\operatorname{SE}}(\widehat{\Delta})=.15.

The estimate and standard error answer different questions: the estimate locates the observed contrast, while the standard error describes how much that estimator would vary across samples under the model.

Constructing a normal interval

Suppose repeated samples make the studentized estimator approximately standard normal under the declared sampling model:

\frac{\widehat{\Delta}-\Delta} {\widehat{\operatorname{SE}}(\widehat{\Delta})} \approx N(0,1).

About 95% of a standard normal distribution lies between -1.96 and 1.96. This gives the event

-1.96 \leq \frac{\widehat{\Delta}-\Delta} {\widehat{\operatorname{SE}}(\widehat{\Delta})} \leq 1.96.

Now multiply all three parts by the positive estimated standard error and then move \widehat{\Delta} across the inequalities. These licensed algebraic moves place the unknown parameter between two random endpoints:

\widehat{\Delta} -1.96\widehat{\operatorname{SE}}(\widehat{\Delta}) \leq \Delta \leq \widehat{\Delta} +1.96\widehat{\operatorname{SE}}(\widehat{\Delta}).

For the observed study,

.42\mathbin{\pm}1.96(.15) =.42\mathbin{\pm}.294.

The realized interval is thus

[.126,.714].

Interpreting coverage across samples

Before data are sampled, the endpoints are random variables L(\mathbf{X}) and U(\mathbf{X}). The population parameter \Delta is fixed. Write \mathbb{P}_{\Delta} for probability under the sampling model indexed by that fixed parameter. The coverage probability of the interval procedure is the probability that the random interval contains the fixed parameter. A 95% confidence procedure satisfies

\mathbb{P}_{\Delta} \left( \left\{\omega\in\Omega\mid L(\mathbf{X}(\omega))\leq\Delta\leq U(\mathbf{X}(\omega)) \right\} \right) \approx.95.

Imagine repeating the complete study many times. Each sample produces a different estimate, standard error, and interval. About 95% of those intervals contain the fixed parameter when the model and approximation are appropriate.

Code
set.seed(118)
delta <- .40
se <- .15
estimate <- rnorm(100000, mean = delta, sd = se)
lower <- estimate - 1.96 * se
upper <- estimate + 1.96 * se
coverage <- mean(lower <= delta & delta <= upper)

stopifnot(abs(coverage - .95) < .01)
coverage

This simulation uses a known standard error and an exactly normal sampling distribution, so its coverage should be close to .95. It illustrates the repeated-sampling definition. It does not verify the assumptions for the prominence study.

Interpreting the realized interval

After observing the data, the endpoints .126 and .714 are fixed. The parameter either lies inside them or it does not. Under the frequentist model used here, the statement “there is a 95% probability that \Delta lies in this interval” is not licensed.

That statement transfers a probability assigned to varying intervals onto a fixed parameter and one fixed interval. A Bayesian credible interval can support a posterior probability statement, but it is a different inferential object. The distinction is developed on the page about posterior summaries.

The interval also does not describe the distribution of participant responses. It does not imply that 95% of individual prominence differences fall between .126 and .714. Response variation and estimator uncertainty are different quantities.

Stating the conditions for coverage

The nominal 95% rate depends on the standard error and the reference distribution. Dependence among observations, severe skew at a small sample size, or a misspecified response model can change the actual coverage. Reporting “95%” does not override those conditions.

Check your understanding

  1. Which object is fixed before sampling, the interval or \Delta?
  2. What changes across repeated studies?
  3. Recalculate the interval if the estimated standard error is .25.
  4. Why does a confidence interval for a mean not contain 95% of the observations?