The MCMC page separated the posterior target from the dependent sequence used to approximate it. A trace plot places iteration on the horizontal axis and the sampled parameter value on the vertical axis, with independently initialized chains drawn as separate paths.
Suppose \Theta is a parameter representing an association between lexical frequency and responses in a word-recognition model. That parameter does not by itself identify a recognition mechanism. A trace plot asks the narrower computational question of whether the chains repeatedly visit the same stable range of values.
Reading a stable-looking trace
A useful trace has three visible properties: first, each chain fluctuates around a stable location rather than drifting across iterations; second, chains initialized at different values overlap and cross; and third, each chain continues to move rather than remaining nearly constant for long stretches.
Together, these patterns suggest that the chains are exploring the same region. They are not a proof that the chains represent the complete posterior.
Constructing contrasting displays
The following code creates four stable-looking sequences and four separated sequences. They are diagnostic illustrations, not output from a posterior sampler.
In the second display, chain identity predicts the parameter value. Combining those chains into one posterior summary would conceal their disagreement.
Inspecting three visible patterns
Different chains may remain in different regions. This separation may reflect disconnected posterior regions, weak identification, or insufficient movement between regions.
A trace that drifts systematically upward or downward is evidence of nonstationarity over the saved iterations. In this context, nonstationarity means that the distribution represented by the chain appears to change with iteration. A visible drift is a warning, not a proof of its unique cause.
A chain may also remain nearly constant for a long run. It may be exploring a target region very slowly, so many saved iterations add little new information.
These visible patterns do not automatically diagnose their causes. A trace plot can indicate where to investigate, but it cannot select a repair by itself.
Understanding initialization and warmup
Chains are commonly initialized from dispersed values. Early transitions may still reflect those initial states while the sampler adapts and moves toward a posterior region. Stan separates this warmup period from saved sampling iterations. During adaptive warmup, the step size and mass matrix change, so these states do not arise from the fixed transition used for retained sampling.
Discarding adaptive warmup is part of Stan’s algorithmic procedure. It is not evidence that every retained chain has reached the target. Discarding arbitrary portions of a problematic saved trace until it looks stable is not a general repair. If saved chains drift or remain separated, the model should be refit after addressing the computational cause.
An overlapping trace is not proof of convergence
An overlapping dense trace is sometimes described as a “hairy caterpillar” and treated as proof of convergence. All chains can overlap in one region while missing another, and dense lines can hide strong dependence between iterations.
Trace plots assess computation. They do not test whether the response family, predictors, or dependence structure provide an adequate model of linguistic observations. A good trace from a poor model is still a good computation of the wrong target.
Check your understanding
What does persistent separation between chains show directly?
Why does a flat segment indicate little new Monte Carlo information?
Why can overlapping traces still miss a posterior region?
Which question requires a posterior predictive check rather than a trace plot?
The \widehat{R} page turns cross-chain agreement into a numerical comparison.
---title: "Reading trace plots"---The [MCMC page](markov-chain-monte-carlo.qmd) separated the posterior target from the dependent sequence used to approximate it. A [**trace plot**](https://mc-stan.org/bayesplot/articles/plotting-mcmc-draws.html) places iteration on the horizontal axis and the sampled parameter value on the vertical axis, with independently initialized chains drawn as separate paths.Suppose $\Theta$ is a parameter representing an association between lexical frequency and responses in a word-recognition model. That parameter does not by itself identify a recognition mechanism. A trace plot asks the narrower computational question of whether the chains repeatedly visit the same stable range of values.## Reading a stable-looking traceA useful trace has three visible properties: first, each chain fluctuates around a stable location rather than drifting across iterations; second, chains initialized at different values overlap and cross; and third, each chain continues to move rather than remaining nearly constant for long stretches.Together, these patterns suggest that the chains are exploring the same region. They are not a proof that the chains represent the complete posterior.## Constructing contrasting displaysThe following code creates four stable-looking sequences and four separated sequences. They are diagnostic illustrations, not output from a posterior sampler.```{r}#| label: contrasting-trace-plots#| echo: trueset.seed(326)iteration_count <-500chain_count <-4stable <-replicate( chain_count,as.numeric(arima.sim(list(ar = .60), n = iteration_count,sd = .20)))separated <-cbind( stable[, 1] -1.5, stable[, 2] -1.5, stable[, 3] +1.5, stable[, 4] +1.5)old_par <-par(mfrow =c(1, 2))matplot(stable, type ="l", lty =1,xlab ="iteration", ylab ="sampled beta",main ="overlapping chains")matplot(separated, type ="l", lty =1,xlab ="iteration", ylab ="sampled beta",main ="separated chains")par(old_par)stopifnot(diff(range(colMeans(stable))) < .2)stopifnot(diff(range(colMeans(separated))) >2)```In the second display, chain identity predicts the parameter value. Combining those chains into one posterior summary would conceal their disagreement.## Inspecting three visible patternsDifferent chains may remain in different regions. This separation may reflect disconnected posterior regions, weak identification, or insufficient movement between regions.A trace that drifts systematically upward or downward is evidence of [**nonstationarity**](https://mc-stan.org/docs/reference-manual/mcmc.html) over the saved iterations. In this context, nonstationarity means that the distribution represented by the chain appears to change with iteration. A visible drift is a warning, not a proof of its unique cause.A chain may also remain nearly constant for a long run. It may be exploring a target region very slowly, so many saved iterations add little new information.These visible patterns do not automatically diagnose their causes. A trace plot can indicate where to investigate, but it cannot select a repair by itself.## Understanding initialization and warmupChains are commonly initialized from dispersed values. Early transitions may still reflect those initial states while the sampler adapts and moves toward a posterior region. Stan separates this warmup period from saved sampling iterations. During adaptive warmup, the step size and mass matrix change, so these states do not arise from the fixed transition used for retained sampling.Discarding adaptive warmup is part of Stan's algorithmic procedure. It is not evidence that every retained chain has reached the target. Discarding arbitrary portions of a problematic saved trace until it looks stable is not a general repair. If saved chains drift or remain separated, the model should be refit after addressing the computational cause.## An overlapping trace is not proof of convergenceAn overlapping dense trace is sometimes described as a “hairy caterpillar” and treated as proof of convergence. All chains can overlap in one region while missing another, and dense lines can hide strong dependence between iterations.Trace plots assess computation. They do not test whether the response family, predictors, or dependence structure provide an adequate model of linguistic observations. A good trace from a poor model is still a good computation of the wrong target.## Check your understanding1. What does persistent separation between chains show directly?2. Why does a flat segment indicate little new Monte Carlo information?3. Why can overlapping traces still miss a posterior region?4. Which question requires a posterior predictive check rather than a trace plot?The [$\widehat{R}$ page](rhat.qmd) turns cross-chain agreement into a numerical comparison.