Code
mu <- c(80, 100, 120)
expected_estimate <- .9 * mu + 10
bias <- expected_estimate - mu
stopifnot(isTRUE(all.equal(bias, c(2, 0, -2))))
data.frame(mu, expected_estimate, bias)An estimator has a sampling distribution with both a spread and a center. The standard error describes the spread, while the bias describes whether the center lands on the intended target, the estimand:
\operatorname{Bias}(\widehat{\Theta}) =\mathbb{E}[\widehat{\Theta}]-\theta.
The expectation averages the estimates that the rule would produce across possible samples generated under the model. Bias thus belongs to an estimator under a sampling model.
Suppose the target is the population mean consonant duration \mu, which we estimate with the sample mean \overline{X}. One observed dataset produces the numerical estimate \overline{x}=82 ms.
| object | notation | role |
|---|---|---|
| estimand | \mu | population quantity being targeted |
| estimator | \overline{X} | rule applied to a random sample |
| estimate | \overline{x}=82 | rule applied to the observed sample |
Only the estimator has bias. The estimand is the target, and the estimate is one realized number.
Let X_1,\ldots,X_n have a common population mean \mu. The sample mean is
\overline{X}=\frac{1}{n}\sum_{i=1}^nX_i.
The linearity of expectation result derived earlier gives
\begin{aligned} \mathbb{E}[\overline{X}] &=\mathbb{E}\left[\frac{1}{n}\sum_{i=1}^nX_i\right]\\ &=\frac{1}{n}\sum_{i=1}^n\mathbb{E}[X_i]\\ &=\frac{1}{n}\sum_{i=1}^n\mu\\ &=\frac{n\mu}{n}\\ &=\mu. \end{aligned}
Thus
\operatorname{Bias}(\overline{X}) =\mathbb{E}[\overline{X}]-\mu =0.
Independence was not required for this expectation calculation. A common mean and linearity were sufficient. Independence matters for the usual standard-error formula, but it does not enter this proof of unbiasedness.
Suppose the population mean is 80 ms and one sample yields \overline{x}=82 ms. The realized estimation error is
\overline{x}-\mu=82-80=2\text{ ms}.
The sample mean remains unbiased. Other possible samples would yield estimates below 80 ms, above 80 ms, and perhaps exactly at 80 ms. Their errors average to zero under the model.
An observed estimate is not biased merely because it missed the target in one sample: a realized error describes one sample, while bias describes the center of a distribution over possible samples.
Now consider a different estimator:
\widetilde{M}=.9\overline{X}+10.
This rule moves every sample mean ten percent of the way toward 100 ms. Its expectation is
\begin{aligned} \mathbb{E}[\widetilde{M}] &=.9\mathbb{E}[\overline{X}]+10\\ &=.9\mu+10. \end{aligned}
Its bias is thus
\begin{aligned} \operatorname{Bias}(\widetilde{M}) &=(.9\mu+10)-\mu\\ &=10-.1\mu. \end{aligned}
If \mu=80 ms, then
\mathbb{E}[\widetilde{M}]=.9(80)+10=82
and the bias is 2 ms. If \mu=120 ms, the expectation is 118 ms and the bias is -2 ms. If \mu=100 ms, the bias is zero.
mu <- c(80, 100, 120)
expected_estimate <- .9 * mu + 10
bias <- expected_estimate - mu
stopifnot(isTRUE(all.equal(bias, c(2, 0, -2))))
data.frame(mu, expected_estimate, bias)Bias is a property of an estimator under a particular sampling model and parameter value. The same estimator may have positive bias at one value, negative bias at another, and zero bias at a third. State the estimator, sampling model, and parameter values when describing its bias.
The transformed rule may vary less across samples because it multiplies \overline{X} by .9. That reduced spread could sometimes compensate for its displacement from the target. Unbiasedness alone thus does not rank all estimation procedures.
The mean-squared-error page combines an estimator’s variance and bias into one measure of expected error.