The preceding page separated an estimator’s bias from its sampling variance. The mean squared error (MSE) combines them by asking how far an estimator falls from its estimand on average:

\operatorname{MSE}(\widehat{\Theta}) =\mathbb{E}\left[(\widehat{\Theta}-\theta)^2\right].

The squared distance makes errors in either direction positive and gives more weight to large errors.

Setting up two estimators

Suppose each of four independently sampled complement-clause tokens either contains an overt that or has a zero complementizer. Write X_i=1 for overt that and X_i=0 for the zero variant. This is a constructed teaching example, and the population probability of the overt variant is

\pi=.60.

The estimand is \pi. If K=\sum_{i=1}^4X_i is the number of overt that tokens, the ordinary sample-proportion estimator is

\widehat{\Pi}=\frac{K}{4}.

Now define a second rule that adds one hypothetical overt token and one hypothetical zero token:

\widetilde{\Pi}=\frac{K+1}{6}.

This second estimator pulls every estimate toward .50. For instance, K=4 gives 1 under the ordinary rule but 5/6 under the pulled rule.

Calculating the ordinary estimator’s MSE

For K\sim\operatorname{Binomial}(4,.60),

\mathbb{E}[K]=4(.60)=2.4

and

\operatorname{Var}(K)=4(.60)(.40)=.96.

The ordinary estimator is unbiased because

\mathbb{E}[\widehat{\Pi}] =\frac{\mathbb{E}[K]}{4} =\frac{2.4}{4} =.60.

Its variance is

\operatorname{Var}(\widehat{\Pi}) =\frac{\operatorname{Var}(K)}{4^2} =\frac{.96}{16} =.06.

Decomposing mean squared error

Add and subtract \mathbb{E}[\widehat{\Theta}] inside the squared error:

\widehat{\Theta}-\theta =\left(\widehat{\Theta}-\mathbb{E}[\widehat{\Theta}]\right) +\left(\mathbb{E}[\widehat{\Theta}]-\theta\right).

After squaring and taking expectations, the cross-product has expectation zero. The centered term \widehat{\Theta}-\mathbb{E}[\widehat{\Theta}] has expectation zero, while the bias term is constant under the sampling model. Thus the cross-product is a constant times zero. This leaves

\operatorname{MSE}(\widehat{\Theta}) =\operatorname{Var}(\widehat{\Theta}) +\operatorname{Bias}(\widehat{\Theta})^2.

Because \widehat{\Pi} is unbiased, its MSE equals its variance:

\operatorname{MSE}(\widehat{\Pi})=.06.

Calculating the pulled estimator’s MSE

The expected value of the pulled estimator is

\mathbb{E}[\widetilde{\Pi}] =\frac{\mathbb{E}[K]+1}{6} =\frac{2.4+1}{6} \approx.5667.

Its bias at \pi=.60 is

.5667-.60\approx-.0333.

Adding one to K does not change its variance. Dividing by six gives

\operatorname{Var}(\widetilde{\Pi}) =\frac{\operatorname{Var}(K)}{6^2} =\frac{.96}{36} \approx.0267.

Thus

\begin{aligned} \operatorname{MSE}(\widetilde{\Pi}) &=.0267+(-.0333)^2\\ &\approx.0278. \end{aligned}

At \pi=.60 and n=4, the biased rule has lower MSE. Its reduction in variance is larger than its squared bias.

Code
k <- 0:4
probability <- dbinom(k, size = 4, prob = .60)
ordinary <- k / 4
pulled <- (k + 1) / 6

mse_ordinary <- sum(probability * (ordinary - .60)^2)
mse_pulled <- sum(probability * (pulled - .60)^2)

stopifnot(isTRUE(all.equal(mse_ordinary, .06)))
stopifnot(isTRUE(all.equal(mse_pulled, 1 / 36)))
c(mse_ordinary = mse_ordinary, mse_pulled = mse_pulled)

The code enumerates every possible count, weights its squared error by its binomial probability, and recovers the analytic results.

Unbiasedness does not rank estimators by itself

An unbiased estimator is not necessarily preferable to every biased estimator, since bias is one component of expected error rather than a complete ordering of procedures. The lower-MSE rule can be biased.

This comparison also depends on \pi and n. A procedure that has lower MSE at one parameter value may have higher MSE elsewhere. The calculation supports a local comparison under the stated model, not a universal claim about the estimator.

Check your understanding

  1. An estimator has variance .04 and bias .10. What is its MSE?
  2. Which term in the decomposition disappears for an unbiased estimator?
  3. Why does adding one to K leave its variance unchanged?
  4. What must be specified before two estimators can be ranked by MSE?

We now return to uncertainty about a parameter. A confidence interval uses a sampling procedure to produce a range of parameter values.