Null hypotheses and test statistics

The sampling distribution page showed why a sample estimate can differ from its estimand. A significance test compares an observed discrepancy with the discrepancies generated by a specified reference model.

The comparison requires three objects: a null hypothesis, a test statistic, and a p value. We introduce them in that order because each one answers a different question: what claim supplies the reference, how far did the estimate fall from it, and how unusual is a discrepancy at least that large under the reference model?

Stating the parameter claim

Suppose a production study measures a duration difference between two consonant contexts. Each independent unit supplies one difference. Let the population mean difference be \mu_D.

The null hypothesis fixes this parameter at a reference value:

H_0:\mu_D=0.

A two-sided alternative hypothesis identifies departures in either direction as relevant:

H_1:\mu_D\neq0.

These hypotheses concern a parameter under a model. They do not divide datasets into boring and interesting cases. The sample contains an estimate, not a true or false null hypothesis.

The reference value should follow from the linguistic question. If the claim to be tested fixes the mean difference at 5 ms, the corresponding point null is \mu_D=5\text{ ms} rather than zero. A conventional reference value cannot substitute for the claim under evaluation.

Standardizing the discrepancy

Let \widehat{\mu}_D be the sample mean difference, let \widehat{\mu}_d be its realized estimate, and let \widehat{\operatorname{SE}}(\widehat{\mu}_D) be its estimated standard error. The statistic

Z = \frac{ \widehat{\mu}_D-\mu_{D,0} }{ \widehat{\operatorname{SE}}(\widehat{\mu}_D) }

expresses the discrepancy from the null value \mu_{D,0} in standard-error units.

Suppose the estimate is 12 ms, the null value is 0 ms, and the estimated standard error is 5 ms. Then

z_{\mathrm{obs}} =\frac{12-0}{5} =2.4.

The numerator records direction and magnitude in milliseconds, while the denominator records sampling uncertainty in the same units, making their ratio unitless. A statistic of 2.4 says that the estimate is 2.4 estimated standard errors above the null value.

The statistic is not a probability. To calculate a probability, we need its null distribution, which is the distribution of statistic values generated when the null hypothesis and the sampling assumptions hold.

Calculating a tail probability

Suppose a large-sample argument supplies an approximately standard normal reference distribution. For a two-sided test, results at least as discrepant as 2.4 lie in both tails:

p = \mathbb{P}_{H_0} \left( |Z|\geq|2.4| \right).

Symmetry lets us calculate one tail and double it:

Code
z_observed <- 2.4
p_value <- 2 * pnorm(-abs(z_observed))

stopifnot(abs(p_value - 0.016395) < 0.000001)
p_value

The p value is approximately .016. Under the null hypothesis and the stated model, a statistic at least this far from zero would occur in about 1.6% of repetitions.

The phrase under the null hypothesis and the stated model scopes the probability. Saying that the null hypothesis has probability .016 reverses the conditional statement. The calculation assumes H_0 and evaluates possible statistics. It does not assign a probability to H_0.

The p value also is not the probability that the result will replicate, the probability that the analysis made an error, or the proportion of observations caused by chance.

Separating effect size from compatibility

The test statistic combines the estimate and its precision: a 2 ms estimate with a standard error of .5 ms produces z=4, while a 20 ms estimate with a standard error of 10 ms produces z=2. The first statistic is larger even though its raw effect is smaller.

Thus a p value cannot communicate linguistic magnitude by itself. Report the estimated contrast in meaningful units and an interval alongside the test: the estimate addresses magnitude, while the p value addresses compatibility with one null value under one reference procedure.

Treating a threshold as part of the procedure

A rule such as p<.05 must be fixed as part of the analysis procedure if it is to have its usual long-run error interpretation. Trying many outcomes, contrasts, or exclusion rules and reporting only the smallest p value changes that procedure.

Passing a threshold does not establish the theory that motivated the contrast. It rules against one parameter value under a model. Competing linguistic explanations may make the same directional prediction.

NoteProblem-set reading

PS1, Task 3: paired-mean inference constructs the estimate, standard error, test statistic, and tail probability from within-speaker first-formant (F_1) differences. The paired-observations page defines this acoustic response. The reading is included because the calculation keeps the null value, standardized discrepancy, and tail probability in distinct roles.

Check your understanding

  1. What population quantity does H_0:\mu_D=0 constrain?
  2. Why is z_{\mathrm{obs}}=2.4 not a probability?
  3. What would the statistic be if the estimate stayed at 12 ms but the standard error increased to 8 ms?
  4. Why does p=.016 not mean that the null hypothesis has probability .016?