An estimator has a sampling distribution with both a spread and a center. The standard error describes the spread, while the bias describes whether the center lands on the intended target, the estimand:

\operatorname{Bias}(\widehat{\Theta}) =\mathbb{E}[\widehat{\Theta}]-\theta.

The expectation averages the estimates that the rule would produce across possible samples generated under the model. Bias thus belongs to an estimator under a sampling model.

Keeping three objects distinct

Suppose the target is the population mean consonant duration \mu, which we estimate with the sample mean \overline{X}. One observed dataset produces the numerical estimate \overline{x}=82 ms.

object notation role
estimand \mu population quantity being targeted
estimator \overline{X} rule applied to a random sample
estimate \overline{x}=82 rule applied to the observed sample

Only the estimator has bias. The estimand is the target, and the estimate is one realized number.

Showing that the sample mean is unbiased

Let X_1,\ldots,X_n have a common population mean \mu. The sample mean is

\overline{X}=\frac{1}{n}\sum_{i=1}^nX_i.

The linearity of expectation result derived earlier gives

\begin{aligned} \mathbb{E}[\overline{X}] &=\mathbb{E}\left[\frac{1}{n}\sum_{i=1}^nX_i\right]\\ &=\frac{1}{n}\sum_{i=1}^n\mathbb{E}[X_i]\\ &=\frac{1}{n}\sum_{i=1}^n\mu\\ &=\frac{n\mu}{n}\\ &=\mu. \end{aligned}

Thus

\operatorname{Bias}(\overline{X}) =\mathbb{E}[\overline{X}]-\mu =0.

Independence was not required for this expectation calculation. A common mean and linearity were sufficient. Independence matters for the usual standard-error formula, but it does not enter this proof of unbiasedness.

Separating realized error from bias

Suppose the population mean is 80 ms and one sample yields \overline{x}=82 ms. The realized estimation error is

\overline{x}-\mu=82-80=2\text{ ms}.

The sample mean remains unbiased. Other possible samples would yield estimates below 80 ms, above 80 ms, and perhaps exactly at 80 ms. Their errors average to zero under the model.

An observed estimate is not biased merely because it missed the target in one sample: a realized error describes one sample, while bias describes the center of a distribution over possible samples.

Constructing a biased rule

Now consider a different estimator:

\widetilde{M}=.9\overline{X}+10.

This rule moves every sample mean ten percent of the way toward 100 ms. Its expectation is

\begin{aligned} \mathbb{E}[\widetilde{M}] &=.9\mathbb{E}[\overline{X}]+10\\ &=.9\mu+10. \end{aligned}

Its bias is thus

\begin{aligned} \operatorname{Bias}(\widetilde{M}) &=(.9\mu+10)-\mu\\ &=10-.1\mu. \end{aligned}

If \mu=80 ms, then

\mathbb{E}[\widetilde{M}]=.9(80)+10=82

and the bias is 2 ms. If \mu=120 ms, the expectation is 118 ms and the bias is -2 ms. If \mu=100 ms, the bias is zero.

Code
mu <- c(80, 100, 120)
expected_estimate <- .9 * mu + 10
bias <- expected_estimate - mu

stopifnot(isTRUE(all.equal(bias, c(2, 0, -2))))
data.frame(mu, expected_estimate, bias)

Bias is a property of an estimator under a particular sampling model and parameter value. The same estimator may have positive bias at one value, negative bias at another, and zero bias at a third. State the estimator, sampling model, and parameter values when describing its bias.

Bias is not the only property

The transformed rule may vary less across samples because it multiplies \overline{X} by .9. That reduced spread could sometimes compensate for its displacement from the target. Unbiasedness alone thus does not rank all estimation procedures.

Check your understanding

  1. An unbiased estimator produces an estimate 5 ms above the target. Is this a contradiction?
  2. What is the bias of \widetilde{M} when \mu=60 ms?
  3. At what value of \mu is \widetilde{M} unbiased?
  4. Why did the proof for \overline{X} not require independence?

The mean-squared-error page combines an estimator’s variance and bias into one measure of expected error.