Code
set.seed(118)
delta <- .40
se <- .15
estimate <- rnorm(100000, mean = delta, sd = se)
lower <- estimate - 1.96 * se
upper <- estimate + 1.96 * se
coverage <- mean(lower <= delta & delta <= upper)
stopifnot(abs(coverage - .95) < .01)
coverageThe preceding pages separated a point estimate from its standard error. A confidence interval is a procedure that combines those two quantities to produce a range of parameter values.
The word procedure matters. Confidence is a long-run property of the rule that generates intervals, not a probability attached to the fixed endpoints observed in one study. We will first derive the rule while the data are still random and only then interpret the interval produced by the observed sample.
Suppose a perception study compares mean prosodic-prominence ratings in two conditions. Prominence concerns the extent to which a word or syllable stands out relative to surrounding material. It is a perceptual and phonological notion that may depend on several cues, not another name for duration or fundamental frequency. The population contrast is
\Delta=\mu_A-\mu_B.
This parameter is the estimand, while the difference between sample means, \widehat{\Delta}, is the estimator. The observed data produce the estimate
\widehat{\delta}=.42
rating points. Its estimated standard error is
\widehat{\operatorname{SE}}(\widehat{\Delta})=.15.
The estimate and standard error answer different questions: the estimate locates the observed contrast, while the standard error describes how much that estimator would vary across samples under the model.
Suppose repeated samples make the studentized estimator approximately standard normal under the declared sampling model:
\frac{\widehat{\Delta}-\Delta} {\widehat{\operatorname{SE}}(\widehat{\Delta})} \approx N(0,1).
About 95% of a standard normal distribution lies between -1.96 and 1.96. This gives the event
-1.96 \leq \frac{\widehat{\Delta}-\Delta} {\widehat{\operatorname{SE}}(\widehat{\Delta})} \leq 1.96.
Now multiply all three parts by the positive estimated standard error and then move \widehat{\Delta} across the inequalities. These licensed algebraic moves place the unknown parameter between two random endpoints:
\widehat{\Delta} -1.96\widehat{\operatorname{SE}}(\widehat{\Delta}) \leq \Delta \leq \widehat{\Delta} +1.96\widehat{\operatorname{SE}}(\widehat{\Delta}).
For the observed study,
.42\mathbin{\pm}1.96(.15) =.42\mathbin{\pm}.294.
The realized interval is thus
[.126,.714].
Before data are sampled, the endpoints are random variables L(\mathbf{X}) and U(\mathbf{X}). The population parameter \Delta is fixed. Write \mathbb{P}_{\Delta} for probability under the sampling model indexed by that fixed parameter. The coverage probability of the interval procedure is the probability that the random interval contains the fixed parameter. A 95% confidence procedure satisfies
\mathbb{P}_{\Delta} \left( \left\{\omega\in\Omega\mid L(\mathbf{X}(\omega))\leq\Delta\leq U(\mathbf{X}(\omega)) \right\} \right) \approx.95.
Imagine repeating the complete study many times. Each sample produces a different estimate, standard error, and interval. About 95% of those intervals contain the fixed parameter when the model and approximation are appropriate.
set.seed(118)
delta <- .40
se <- .15
estimate <- rnorm(100000, mean = delta, sd = se)
lower <- estimate - 1.96 * se
upper <- estimate + 1.96 * se
coverage <- mean(lower <= delta & delta <= upper)
stopifnot(abs(coverage - .95) < .01)
coverageThis simulation uses a known standard error and an exactly normal sampling distribution, so its coverage should be close to .95. It illustrates the repeated-sampling definition. It does not verify the assumptions for the prominence study.
After observing the data, the endpoints .126 and .714 are fixed. The parameter either lies inside them or it does not. Under the frequentist model used here, the statement “there is a 95% probability that \Delta lies in this interval” is not licensed.
That statement transfers a probability assigned to varying intervals onto a fixed parameter and one fixed interval. A Bayesian credible interval can support a posterior probability statement, but it is a different inferential object. The distinction is developed on the page about posterior summaries.
The interval also does not describe the distribution of participant responses. It does not imply that 95% of individual prominence differences fall between .126 and .714. Response variation and estimator uncertainty are different quantities.
The nominal 95% rate depends on the standard error and the reference distribution. Dependence among observations, severe skew at a small sample size, or a misspecified response model can change the actual coverage. Reporting “95%” does not override those conditions.