Effective sample size

The autocorrelation page showed why dependent draws contain less information about some summaries than the same number of independent draws. For the basic calculation, suppose one stationary chain contains

S=3{,}000

saved draws for a difference in log fundamental frequency between two phonological tone categories. Fundamental frequency is the acoustic rate of vocal-fold vibration and may cue tone, but the measurement is not itself a tone category. If those draws are strongly dependent, they do not estimate posterior summaries as precisely as 3,000 independent target draws. Modern software extends the calculation across multiple chains.

The question is how many independent target draws would provide comparable precision for a particular summary. The effective sample size (ESS) answers that question approximately.

Connecting ESS to autocorrelation

For a stationary scalar sequence and a particular scalar function g(\Theta), a basic expression is

\operatorname{ESS} \approx \frac{S} {1+2\sum_{t=1}^{\infty}\rho_t},

where \rho_t is the autocorrelation at lag t for the sequence g(\Theta_1),\ldots,g(\Theta_S). The denominator is the integrated autocorrelation time, which positive dependence increases, thereby reducing ESS. Negative dependence can make ESS larger than the raw draw count for some summaries, so ESS is not literally a subset of retained iterations.

Calculating a geometric case

Suppose

\rho_t=.5^t.

The sum of this geometric series is

\sum_{t=1}^{\infty}.5^t =1.

Thus

1+2\sum_{t=1}^{\infty}\rho_t =1+2(1) =3.

The effective sample size is

\operatorname{ESS} \approx \frac{3000}{3} =1000.

The chain output has roughly the mean-estimation precision of 1,000 independent target draws.

Code
saved_draws <- 3000
rho_sum <- sum(.5^(1:100))
integrated_time <- 1 + 2 * rho_sum
ess <- saved_draws / integrated_time

stopifnot(abs(rho_sum - 1) < 1e-12)
stopifnot(abs(ess - 1000) < 1e-8)
c(raw_draws = saved_draws,
  integrated_time = integrated_time,
  ESS = ess)

Connecting ESS to Monte Carlo error

If g(\Theta) has posterior standard deviation \sigma_g, the Monte Carlo standard error of its estimated mean is approximately

\operatorname{MCSE} \approx \frac{\sigma_g}{\sqrt{\operatorname{ESS}}}.

Suppose the posterior standard deviation of the tone contrast is .30. With ESS equal to 1,000,

\operatorname{MCSE} \approx \frac{.30}{\sqrt{1000}} \approx.0095.

This is numerical uncertainty in the estimated posterior mean. It is not the posterior standard deviation .30 and does not describe uncertainty about the linguistic parameter.

Distinguishing bulk and tail information

Bulk ESS is the ESS of rank-normalized draws. It is a robust diagnostic for location summaries such as means and medians, though it is not exactly the ESS of the raw-parameter mean. Tail ESS is the smaller of ESS estimates associated with the pooled 5% and 95% quantiles, so it diagnoses exploration relevant to interval endpoints and tail probabilities.

A chain may move adequately through the center while rarely reaching a posterior tail. Its mean can then be estimated precisely while a 2.5% quantile remains unstable. The required ESS depends on the quantity being reported.

Modern ESS estimates combine multiple chains and stabilize the estimated autocorrelation sequence before summing it. The infinite-series formula is the conceptual target, while software must estimate it from finite draws. Current Stan guidance recommends bulk and tail ESS of roughly 100 per chain for final summaries, but the inferential requirement is that MCSE be small enough for the precision of the reported claim. The count threshold supports reliable diagnostics; it is not a universal guarantee of substantive precision.

Effective sample size does not count participants

ESS counts simulation information, not participants, speakers, items, or tokens. An ESS of 1,000 thus does not mean that the study has 1,000 observations.

Running longer can increase ESS and reduce Monte Carlo error. It cannot add information to the posterior target. Low ESS calls the numerical summary into question, while wide posterior uncertainty may remain even after ESS is excellent.

Raw draw count is thus a storage fact, while ESS is a summary-specific estimate of computational information. Neither is a measure of linguistic sample size.

NoteProblem Set 4 connection

Problem Set 4, Task 3 reports both bulk and tail effective sample sizes for the model quantities used in the projection analysis. A stable posterior center does not establish that interval endpoints are estimated with comparable simulation information, so both summaries are required.

Check your understanding

  1. Why does positive autocorrelation usually reduce ESS?
  2. If the integrated autocorrelation time is five and S=10{,}000, what is ESS?
  3. Why can bulk ESS and tail ESS differ?
  4. Which quantity changes when the sampler runs longer, posterior spread or Monte Carlo error?