Sampling distributions

The maximum likelihood page produced one estimate from one observed sample. We now study an estimator as a rule that could be applied to every possible sample. In a lexical decision task, a participant classifies each presented letter string as a word or a nonword. Suppose the population probability of a correct classification is

\pi=.75.

We sample 20 independent trials and estimate that accuracy probability with the sample proportion, which might be .70 in one sample, .80 in another, and .75 in a third. The population parameter is unchanged while the sample summary varies. For current purposes, this is a teaching model: real lexical decision trials may share participants and items, so trial-level independence would need justification in an actual experiment.

Basically, we apply the same rule to every sample the model could have produced and collect the resulting estimates. More specifically, the distribution of that estimation rule across possible samples is its sampling distribution.

Keeping three objects distinct

The estimand is the target population quantity:

\pi=.75.

The estimator is the rule applied before we know which sample will occur:

\widehat{\Pi} =\frac{1}{20}\sum_{i=1}^{20}X_i.

The capital notation emphasizes that \widehat{\Pi} varies across possible samples.

After one sample is observed, the estimator produces an estimate, perhaps

\widehat{\pi}=.70.

object varies across repeated samples? role
estimand \pi no fixed target under the model
estimator \widehat{\Pi} yes rule applied to any possible sample
estimate \widehat{\pi} already realized output for one observed sample

An estimate is one realized number. Bias, variance, and standard error are properties of the estimator across possible samples, so the distinction affects every calculation that follows.

Enumerating a very small sampling distribution

Before simulating 20 trials, take n=2 and \pi=.75. The possible ordered samples and proportion estimates are

sample probability estimate
(0,0) .25^2=.0625 0
(0,1) .25(.75)=.1875 .5
(1,0) .75(.25)=.1875 .5
(1,1) .75^2=.5625 1

Combine samples that produce the same estimate:

possible \widehat{\Pi} probability
0 .0625
.5 .375
1 .5625

This PMF is the sampling distribution for the sample proportion when n=2 and \pi=.75.

Using the expectation rule developed earlier, its expected value is

0(.0625)+.5(.375)+1(.5625)=.75.

The center equals the estimand.

Simulating the larger case

For n=20, enumeration is possible but tedious. Simulation gives an approximation.

Code
set.seed(414)

accuracy_estimates <- replicate(
  10000,
  mean(rbinom(20, size = 1, prob = .75))
)

mean(accuracy_estimates)
sd(accuracy_estimates)

The simulated center should be near .75. The simulated standard deviation should be near the analytic value derived below.

Calculating center and spread

Under IID Bernoulli sampling,

\mathbb{E}[\widehat{\Pi}]=\pi

and

\operatorname{Var}(\widehat{\Pi}) =\frac{\pi(1-\pi)}{n}.

For \pi=.75 and n=20,

\operatorname{Var}(\widehat{\Pi}) =\frac{.75(.25)}{20} =.009375.

The sampling standard deviation is

\sqrt{.009375}\approx.0968.

Estimates from samples of 20 can thus differ from the target by several percentage points under the model.

Changing the sample size

If n increases to 100, the variance becomes

\frac{.75(.25)}{100}=.001875,

and the sampling standard deviation becomes approximately .0433. Though the population parameter has not changed, the estimator has become more stable because each estimate averages more independent observations.

A sampling distribution contains estimates

The sampling distribution here contains possible sample proportions, not zero and one trial responses. It is a distribution of estimates across possible samples, not a distribution of individual response values.

State what repeats. Entire samples repeat, and one estimate is computed from each sample.

Check your understanding

  1. For n=2 and \pi=.50, enumerate the sampling distribution of the sample proportion.
  2. Which object is fixed in the repeated-sampling model, \pi or \widehat{\Pi}?
  3. Why does increasing n narrow the sampling distribution without changing the population accuracy?
  4. Distinguish an estimator from its realized estimate in one sentence.

The next pages use this distribution in two ways. Standard errors summarize its spread, while estimator bias compares its center with the target.