Suppose a participant hears The boys carried a piano up the stairs and chooses between a picture in which the boys act together and one in which each boy acts separately. These are the collective and distributive readings of a plural predication. They are not reciprocal readings: a reciprocal would express a relation that the boys bear to one another. Let X_i=1 mark a collective response on trial i, and let \Theta be the probability of that response in the declared population and task.
Bayesian inference represents uncertainty about \Theta with a probability distribution, whose form before observing the present responses is the prior distribution. The question is how the response data redistribute that probability. This redistributed probability forms the posterior distribution.
Beginning with a finite parameter space
To make the update visible by hand, suppose the model initially permits three possible response probabilities:
\Theta\in\{.25,.50,.75\}.
Assign prior probabilities
\theta
\mathbb{P}(\Theta=\theta)
.25
.20
.50
.50
.75
.30
These numbers describe uncertainty about the parameter before the present sample; they do not say that a response is 20%, 50%, or 30% likely to be collective. Conditional on \Theta=\theta, the Bernoulli model supplies the response probability \theta.
Now observe the ordered response vector \mathbf{x}=(1,1,1,0). Conditional on \Theta=\theta, assume the four responses are independent Bernoulli variables with a shared response probability. The probability mass of this particular vector is
For fixed \theta, this expression is one value in the probability distribution over possible response vectors. Once \mathbf{x} is observed, the same expression viewed as a function of \theta is the likelihood. It is not a probability distribution over the three parameter values.
For the finite parameter space, we can update in three steps: first, evaluate the likelihood at each value; second, multiply by the corresponding prior probability; and third, divide each product by the sum of all three products.
\theta
data probability, viewed as likelihood
prior probability
unnormalized posterior
.25
.01171875
.20
.00234375
.50
.0625
.50
.03125
.75
.10546875
.30
.031640625
The normalizing constant is
Z
=.00234375+.03125+.031640625
=.065234375.
Dividing each product by Z gives posterior probabilities of approximately .0359, .4791, and .4850.
The data favor larger values because three of four responses were collective. The prior gives substantial mass to .50, so the posterior probabilities for .50 and .75 are nearly equal. This result is an update from a particular prior through a particular likelihood.
is the prior predictive probability of the particular observed vector. It does not depend on which \theta is being evaluated. In the current calculation it equals Z=.065234375, and dividing by it makes the posterior probabilities sum to one.
For a continuous parameter, prior and posterior uncertainty are represented by densities. With the same discrete response vector, the normalizer becomes
when we first care about posterior shape. Here p_{\mathbf X\mid\Theta} is a PMF over the discrete response vector and f_\Theta is a density over the continuous parameter. A normalized posterior density is still required before assigning posterior probabilities or calculating expectations.
Interpreting a continuous density
For a continuous parameter, f_{\Theta\mid\mathbf{X}}(\theta\mid\mathbf{x}) is a density rather than the probability that \Theta equals exactly \theta. The probability of any single point is zero. Posterior probability is assigned to an interval by integrating:
A higher density at .75 than at .50 says that a small neighborhood around .75 receives more posterior mass per unit width. It does not assign positive probability to the exact point .75 in a continuous model.
Keeping the conditioning direction visible
The error of transposing a conditional probability also applies here. The likelihood evaluates the observed data as a function of a candidate parameter, whereas the posterior distributes uncertainty over parameter values after conditioning on those data. Bayes’ rule connects the two quantities, but it does not make them interchangeable.
Posterior uncertainty is also not computational error. In the hand calculation, the posterior probabilities are exact under the model. Later simulation methods will approximate posterior summaries and add Monte Carlo error, but that error is not part of the posterior itself.
Check your understanding
Which quantity is fixed when evaluating p_{\mathbf X\mid\Theta}(\mathbf{x}\mid\theta)?
Why do the unnormalized posterior values not sum to one?
What role does p_{\mathbf X}(\mathbf{x}) play in Bayes’ rule?
If the prior put all its probability on \theta=.25, could these data create posterior support at .75 under this model?
Why is a continuous posterior density at one point not a point probability?
The posterior-summaries page shows how a posterior distribution supports probability statements, expectations, and intervals.
---title: "Posterior distributions"---The foundations chapter derived [Bayes' rule](../foundations/bayes-rule.qmd), and the frequentist sequence distinguished a [likelihood from a probability distribution over observations](maximum-likelihood.qmd). We now use both distinctions to construct a posterior distribution.Suppose a participant hears *The boys carried a piano up the stairs* and chooses between a picture in which the boys act together and one in which each boy acts separately. These are the [collective and distributive readings of a plural predication](https://pmc.ncbi.nlm.nih.gov/articles/PMC2864947/). They are not reciprocal readings: a reciprocal would express a relation that the boys bear to one another. Let $X_i=1$ mark a collective response on trial $i$, and let $\Theta$ be the probability of that response in the declared population and task.Bayesian inference represents uncertainty about $\Theta$ with a probability distribution, whose form before observing the present responses is the [**prior distribution**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html). The question is how the response data redistribute that probability. This redistributed probability forms the [**posterior distribution**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html).## Beginning with a finite parameter spaceTo make the update visible by hand, suppose the model initially permits three possible response probabilities:$$\Theta\in\{.25,.50,.75\}.$$Assign prior probabilities| $\theta$ | $\mathbb{P}(\Theta=\theta)$ ||---:|---:|| $.25$ | $.20$ || $.50$ | $.50$ || $.75$ | $.30$ |These numbers describe uncertainty about the parameter before the present sample; they do not say that a response is 20%, 50%, or 30% likely to be collective. Conditional on $\Theta=\theta$, the Bernoulli model supplies the response probability $\theta$.Now observe the ordered response vector $\mathbf{x}=(1,1,1,0)$. Conditional on $\Theta=\theta$, assume the four responses are independent Bernoulli variables with a shared response probability. The probability mass of this particular vector is$$\mathbb{P}(\mathbf{X}=\mathbf{x}\mid\Theta=\theta)=\theta^3(1-\theta).$$For fixed $\theta$, this expression is one value in the probability distribution over possible response vectors. Once $\mathbf{x}$ is observed, the same expression viewed as a function of $\theta$ is the likelihood. It is not a probability distribution over the three parameter values.## Multiplying prior support by data supportBayes' rule is$$\mathbb{P}(\Theta=\theta\mid\mathbf{X}=\mathbf{x})=\frac{\mathbb{P}(\mathbf{X}=\mathbf{x}\mid\Theta=\theta)\mathbb{P}(\Theta=\theta)}{\mathbb{P}(\mathbf{X}=\mathbf{x})}.$$For the finite parameter space, we can update in three steps: first, evaluate the likelihood at each value; second, multiply by the corresponding prior probability; and third, divide each product by the sum of all three products.| $\theta$ | data probability, viewed as likelihood | prior probability | unnormalized posterior ||---:|---:|---:|---:|| $.25$ | $.01171875$ | $.20$ | $.00234375$ || $.50$ | $.0625$ | $.50$ | $.03125$ || $.75$ | $.10546875$ | $.30$ | $.031640625$ |The normalizing constant is$$Z=.00234375+.03125+.031640625=.065234375.$$Dividing each product by $Z$ gives posterior probabilities of approximately $.0359$, $.4791$, and $.4850$.```{r}#| label: discrete-posterior-update#| echo: truetheta <-c(.25, .50, .75)prior <-c(.20, .50, .30)likelihood_kernel <- theta^3* (1- theta)unnormalized <- prior * likelihood_kernelposterior <- unnormalized /sum(unnormalized)stopifnot(abs(sum(posterior) -1) <1e-12)stopifnot(isTRUE(all.equal( posterior,c(0.03592814, 0.47904192, 0.48502994),tolerance =1e-7)))data.frame(theta, prior, likelihood_kernel, posterior)```The data favor larger values because three of four responses were collective. The prior gives substantial mass to $.50$, so the posterior probabilities for $.50$ and $.75$ are nearly equal. This result is an update from a particular prior through a particular likelihood.## Understanding the normalizerThe denominator$$\mathbb{P}(\mathbf{X}=\mathbf{x})=\sum_{\theta\in\operatorname{cod}(\Theta)}\mathbb{P}(\mathbf{X}=\mathbf{x}\mid\Theta=\theta)\mathbb{P}(\Theta=\theta)$$is the prior predictive probability of the particular observed vector. It does not depend on which $\theta$ is being evaluated. In the current calculation it equals $Z=.065234375$, and dividing by it makes the posterior probabilities sum to one.For a continuous parameter, prior and posterior uncertainty are represented by densities. With the same discrete response vector, the normalizer becomes$$\mathbb{P}(\mathbf{X}=\mathbf{x})=\int\mathbb{P}(\mathbf{X}=\mathbf{x}\mid\Theta=\theta)f_\Theta(\theta)\,\mathrm{d}\theta.$$The posterior density is thus$$f_{\Theta\mid\mathbf{X}}(\theta\mid\mathbf{x})=\frac{\mathbb{P}(\mathbf{X}=\mathbf{x}\mid\Theta=\theta)f_\Theta(\theta)}{\mathbb{P}(\mathbf{X}=\mathbf{x})}.$$We retain the distinction between the observation PMF and the parameter density. The shorter proportionality is$$f_{\Theta\mid\mathbf X}(\theta\mid\mathbf{x})\proptop_{\mathbf X\mid\Theta}(\mathbf{x}\mid\theta)f_\Theta(\theta)$$when we first care about posterior shape. Here $p_{\mathbf X\mid\Theta}$ is a PMF over the discrete response vector and $f_\Theta$ is a density over the continuous parameter. A normalized posterior density is still required before assigning posterior probabilities or calculating expectations.## Interpreting a continuous densityFor a continuous parameter, $f_{\Theta\mid\mathbf{X}}(\theta\mid\mathbf{x})$ is a density rather than the probability that $\Theta$ equals exactly $\theta$. The probability of any single point is zero. Posterior probability is assigned to an interval by integrating:$$\mathbb{P}(a<\Theta<b\mid\mathbf{X}=\mathbf{x})=\int_a^bf_{\Theta\mid\mathbf{X}}(\theta\mid\mathbf{x})\,\mathrm{d}\theta.$$A higher density at $.75$ than at $.50$ says that a small neighborhood around $.75$ receives more posterior mass per unit width. It does not assign positive probability to the exact point $.75$ in a continuous model.## Keeping the conditioning direction visibleThe [error of transposing a conditional probability](../foundations/bayes-rule.qmd#the-order-of-the-events) also applies here. The likelihood evaluates the observed data as a function of a candidate parameter, whereas the posterior distributes uncertainty over parameter values after conditioning on those data. Bayes' rule connects the two quantities, but it does not make them interchangeable.Posterior uncertainty is also not computational error. In the hand calculation, the posterior probabilities are exact under the model. Later simulation methods will approximate posterior summaries and add Monte Carlo error, but that error is not part of the posterior itself.## Check your understanding1. Which quantity is fixed when evaluating $p_{\mathbf X\mid\Theta}(\mathbf{x}\mid\theta)$?2. Why do the unnormalized posterior values not sum to one?3. What role does $p_{\mathbf X}(\mathbf{x})$ play in Bayes' rule?4. If the prior put all its probability on $\theta=.25$, could these data create posterior support at $.75$ under this model?5. Why is a continuous posterior density at one point not a point probability?The [posterior-summaries page](posterior-summaries.qmd) shows how a posterior distribution supports probability statements, expectations, and intervals.