The conjugate prior page updated a prior after observing data. A prior predictive distribution instead propagates the prior through the observation model before the present data are observed.
Suppose a picture-matching task presents one participant with 10 expressions like The spy saw the cop with binoculars. The prepositional phrase with binoculars can modify the verb phrase, so that the spy uses the binoculars, or the noun phrase the cop, so that the cop has them. These are conventionally called high or verb-phrase attachment and low or noun-phrase attachment. Let \widetilde{K} be the number of verb-phrase attachment pictures selected in a possible 10-item dataset generated before the present participant’s responses are observed.
We propose a prior on the participant’s probability of selecting the verb-phrase attachment picture across the declared item set:
\Pi\sim\operatorname{Beta}(2,2).
Conditional on \Pi=\pi, the model treats the ten binary choices as independent with a shared probability and thus defines a binomial count distribution. This assumption brackets item differences and trial-to-trial dependence. Before observing the participant, what values of \widetilde{K} does the complete model expect?
The prior predictive distribution answers this question. Basically, it asks what the complete model says we might record before the current responses are used. More specifically, it averages possible observations over parameter values under the prior, as also described in the Stan User’s Guide:
The tilde marks \widetilde{K} as a possible future count generated from the model rather than a count already observed. Because \widetilde{K} is discrete, the first factor inside the integral is a probability mass. Because \Pi is continuous, f_\Pi(\pi) is a density.
The expression combines two claims. The binomial observation model says how a fixed value of \pi generates a count. The beta prior says which values of \pi receive substantial probability before the present participant’s data are observed. The integral propagates the second claim through the first.
We can read the integral as a weighted average. For each possible \pi, compute the binomial probability of \widetilde{y}. Then weight that probability by the prior density at \pi and add across the unit interval. Values of \pi that are more plausible under the prior contribute more to the resulting predictive probability.
Simulating one possible dataset
First draw a probability from the prior:
\pi^{(s)}\sim\operatorname{Beta}(2,2).
Then draw ten choices using that probability and count the successes:
Repeat the two steps to approximate the full prior predictive distribution.
One simulation draw makes the averaging concrete. Suppose the prior draw is \pi^{(1)}=.68. We keep that value fixed while generating all ten responses in prior predictive dataset 1. If the simulated count is \widetilde{k}^{(1)}=8, the pair (.68,8) is one draw from the joint prior model. A second prior predictive dataset receives a new draw from the prior because the integral averages predictions over our prior uncertainty about \Pi.
The beta density describes an abstract probability parameter, while the simulated values are counts that the experiment could record. They are often easier to evaluate.
If the prior and observation model predict that nearly every ten-trial dataset must contain either 0 or 10 verb-phrase attachment responses, we should ask whether such nearly categorical responding is plausible for this item set and task. The prior predictive distribution makes that commitment visible before the present responses are used.
The opposite failure can occur as well. A prior concentrated tightly around \pi=.50 may imply that almost every ten-trial count lies near 5. Such a model would rule out a strong preference for either interpretation before the experiment begins. If the linguistic theory permits the modeled response probability to lie far from .50, that prior is too restrictive even if its center appears neutral.
The quantiles describe the central spread of possible recorded counts. The boundary proportion describes how often the model expects categorical response patterns. Neither quantity is a universal diagnostic. We choose summaries that express the parts of the measurement we care about.
Comparing priors through predictions
Parameter-scale language can conceal how strongly two priors differ. Consider three beta priors:
All three are symmetric around .50. But symmetry does not make them equivalent. The \operatorname{Beta}(2,2) density has its maximum at .50, the \operatorname{Beta}(20,20) density is much more concentrated there, and the \operatorname{Beta}(.5,.5) density increases toward both boundaries.
The predictive counts state the practical commitments of each prior in the units the experiment records. We can ask whether each commitment is compatible with the design, earlier studies, and the range of behavior the theory leaves open.
Matching the simulation to the sampling structure
The current simulation draws one \pi per complete ten-trial dataset and then ten choices conditional on that value. This represents prior uncertainty about one participant-level probability in the simplified model, but it does not define a population distribution of participant preferences.
To represent several participants with different underlying probabilities, we would need an additional population distribution and one participant probability drawn from it for each participant. That hierarchical structure is not part of the current model. Drawing a new \pi before every trial would instead put the beta variation at the trial level. These are different generative claims, even if every version uses the same beta and Bernoulli functions.
A prior predictive check is only as informative as this generative structure. Simulating the correct response range while ignoring the grouping in the study can make a poorly specified model appear plausible.
Keeping prior checking prior to the update
A prior predictive check examines simulated observations from the prior and observation model. It does not use the posterior.
If we repeatedly tune the prior until it reproduces the present observed outcome, that outcome has already influenced the prior. Any such data-dependent choice should be acknowledged rather than described as information supplied before the data.
Prior predictive checking does not require a prior to predict the observed data closely. Before observing the current sample, a useful prior may allow many outcomes that did not happen. The narrower question is whether it places substantial probability on impossible or theoretically indefensible data, or rules out outcomes that the design and existing knowledge make plausible.
NoteProblem Set 4 connection
Problem Set 4, Task 2 uses prior predictive category proportions for a seven category commitment response. Students must thus simulate the response scale before any observed commitment rating updates the model.
NoteCheck the simulation order
Why must each predictive count use the same sampled \pi^{(s)} for all ten choices?
What different model would result from drawing a new \pi before every choice?
Two priors have the same mean. Explain why their prior predictive distributions can still differ.
Give one response scale summary that would detect a prior predicting too many categorical response patterns.
The prior predictive distribution describes possibilities before the present responses update the model. The next page conditions on those responses and predicts a future block.
---title: "Prior predictive distributions"---The [conjugate prior page](conjugate-priors.qmd) updated a prior after observing data. A prior predictive distribution instead propagates the prior through the observation model before the present data are observed.Suppose a picture-matching task presents one participant with 10 expressions like *The spy saw the cop with binoculars*. The prepositional phrase *with binoculars* can modify the verb phrase, so that the spy uses the binoculars, or the noun phrase *the cop*, so that the cop has them. These are conventionally called [high or verb-phrase attachment and low or noun-phrase attachment](https://labs.psychology.illinois.edu/~jkbock/otherpubs/Branigan%20Pickering%20McLean%202005.pdf). Let $\widetilde{K}$ be the number of verb-phrase attachment pictures selected in a possible 10-item dataset generated before the present participant's responses are observed.We propose a prior on the participant's probability of selecting the verb-phrase attachment picture across the declared item set:$$\Pi\sim\operatorname{Beta}(2,2).$$Conditional on $\Pi=\pi$, the model treats the ten binary choices as independent with a shared probability and thus defines a binomial count distribution. This assumption brackets item differences and trial-to-trial dependence. Before observing the participant, what values of $\widetilde{K}$ does the complete model expect?The [**prior predictive distribution**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html) answers this question. Basically, it asks what the complete model says we might record before the current responses are used. More specifically, it averages possible observations over parameter values under the prior, as also described in the [Stan User's Guide](https://mc-stan.org/docs/stan-users-guide/posterior-predictive-checks.html#prior-predictive-model):$$\mathbb{P}(\widetilde{K}=\widetilde{k})=\int\mathbb{P}(\widetilde{K}=\widetilde{k}\mid\Pi=\pi)f_\Pi(\pi)\,\mathrm{d}\pi.$$The tilde marks $\widetilde{K}$ as a possible future count generated from the model rather than a count already observed. Because $\widetilde{K}$ is discrete, the first factor inside the integral is a probability mass. Because $\Pi$ is continuous, $f_\Pi(\pi)$ is a density.The expression combines two claims. The binomial observation model says how afixed value of $\pi$ generates a count. The beta prior says which values of$\pi$ receive substantial probability before the present participant's data areobserved. The integral propagates the second claim through the first.We can read the integral as a weighted average. For each possible $\pi$, computethe binomial probability of $\widetilde{y}$. Then weight that probability by theprior density at $\pi$ and add across the unit interval. Values of $\pi$ that aremore plausible under the prior contribute more to the resulting predictiveprobability.## Simulating one possible datasetFirst draw a probability from the prior:$$\pi^{(s)}\sim\operatorname{Beta}(2,2).$$Then draw ten choices using that probability and count the successes:$$\widetilde{k}^{(s)}\sim\operatorname{Binomial}(10,\pi^{(s)}).$$Repeat the two steps to approximate the full prior predictive distribution.One simulation draw makes the averaging concrete. Suppose the prior draw is $\pi^{(1)}=.68$. We keep that value fixed while generating all ten responses in prior predictive dataset 1. If the simulated count is $\widetilde{k}^{(1)}=8$, the pair $(.68,8)$ is one draw from the joint prior model. A second prior predictive dataset receives a new draw from the prior because the integral averages predictions over our prior uncertainty about $\Pi$.```{r}#| label: prior-predictive-ambiguity#| echo: trueset.seed(414)pi_prior <-rbeta(10000, shape1 =2, shape2 =2)k_prior <-rbinom(10000, size =10, prob = pi_prior)plot(prop.table(table(k_prior)),xlab ="verb-phrase attachment choices out of 10",ylab ="prior predictive probability")```## Inspecting predictions on the response scaleThe beta density describes an abstract probability parameter, while the simulated values are counts that the experiment could record. They are often easier to evaluate.If the prior and observation model predict that nearly every ten-trial dataset must contain either 0 or 10 verb-phrase attachment responses, we should ask whether such nearly categorical responding is plausible for this item set and task. The prior predictive distribution makes that commitment visible before the present responses are used.The opposite failure can occur as well. A prior concentrated tightly around$\pi=.50$ may imply that almost every ten-trial count lies near 5. Such a modelwould rule out a strong preference for either interpretation before the experiment begins. If the linguistic theory permits the modeled response probability to lie far from $.50$, that prior is too restrictive even if its center appears neutral.We can also inspect summaries:```{r}#| label: summarize-prior-predictive#| echo: truequantile(k_prior, c(.025, .25, .5, .75, .975))mean(k_prior ==0| k_prior ==10)```The quantiles describe the central spread of possible recorded counts. Theboundary proportion describes how often the model expects categorical responsepatterns. Neither quantity is a universal diagnostic. We choose summaries thatexpress the parts of the measurement we care about.## Comparing priors through predictionsParameter-scale language can conceal how strongly two priors differ. Considerthree beta priors:$$\operatorname{Beta}(2,2),\qquad\operatorname{Beta}(20,20),\qquad\operatorname{Beta}(.5,.5).$$All three are symmetric around $.50$. But symmetry does not make themequivalent. The $\operatorname{Beta}(2,2)$ density has its maximum at $.50$,the $\operatorname{Beta}(20,20)$ density is much more concentrated there, andthe $\operatorname{Beta}(.5,.5)$ density increases toward both boundaries.```{r}#| label: compare-prior-predictive-counts#| echo: trueset.seed(414)simulate_prior_counts <-function(alpha, beta, draws =10000) { pi_draw <-rbeta(draws, alpha, beta)rbinom(draws, size =10, prob = pi_draw)}k_moderate <-simulate_prior_counts(2, 2)k_concentrated <-simulate_prior_counts(20, 20)k_boundary <-simulate_prior_counts(.5, .5)rbind(moderate =quantile(k_moderate, c(.05, .50, .95)),concentrated =quantile(k_concentrated, c(.05, .50, .95)),boundary =quantile(k_boundary, c(.05, .50, .95)))```The predictive counts state the practical commitments of each prior in the unitsthe experiment records. We can ask whether each commitment is compatible withthe design, earlier studies, and the range of behavior the theory leaves open.## Matching the simulation to the sampling structureThe current simulation draws one $\pi$ per complete ten-trial dataset and then ten choices conditional on that value. This represents prior uncertainty about one participant-level probability in the simplified model, but it does not define a population distribution of participant preferences.To represent several participants with different underlying probabilities, we would need an additional population distribution and one participant probability drawn from it for each participant. That hierarchical structure is not part of the current model. Drawing a new $\pi$ before every trial would instead put the beta variation at the trial level. These are different generative claims, even if every version uses the same beta and Bernoulli functions.A prior predictive check is only as informative as this generative structure.Simulating the correct response range while ignoring the grouping in the studycan make a poorly specified model appear plausible.## Keeping prior checking prior to the updateA [prior predictive check](https://doi.org/10.1111/rssa.12378) examines simulated observations from the prior and observation model. It does not use the posterior.If we repeatedly tune the prior until it reproduces the present observed outcome, that outcome has already influenced the prior. Any such data-dependent choice should be acknowledged rather than described as information supplied before the data.Prior predictive checking does not require a prior to predict the observed dataclosely. Before observing the current sample, a useful prior may allow manyoutcomes that did not happen. The narrower question is whether it placessubstantial probability on impossible or theoretically indefensible data, orrules out outcomes that the design and existing knowledge make plausible.::: {.callout-note title="Problem Set 4 connection"}[Problem Set 4, Task 2](../problem-sets/ps4/ps4.qmd#task-2-check-the-priors-before-fitting-the-data) uses prior predictive category proportions for a seven category commitment response. Students must thus simulate the response scale before any observed commitment rating updates the model.:::::: {.callout-note title="Check the simulation order"}1. Why must each predictive count use the same sampled $\pi^{(s)}$ for all ten choices?2. What different model would result from drawing a new $\pi$ before every choice?3. Two priors have the same mean. Explain why their prior predictive distributions can still differ.4. Give one response scale summary that would detect a prior predicting too many categorical response patterns.:::The prior predictive distribution describes possibilities before the present responses update the model. The [next page](predictive-distributions.qmd) conditions on those responses and predicts a future block.