The conjugate prior page identified a normalized posterior by matching its algebraic form to a known distribution, but many useful prior and likelihood combinations do not supply that match. What, then, remains calculable?
Let the nonnegative product of likelihood and prior have normalizing integral Z. A proper posterior is one that can be normalized to integrate to one, which occurs when 0<Z<\infty. The product need not be positive at every parameter value, and a proper posterior can exist even when Z has no closed-form expression.
Constructing a nonconjugate model
Suppose each of four speakers produces stops in two matched conditions. Closure duration is the interval from the formation of a stop closure to its release, a measure used in phonetic work on distinctions among stops since at least Lisker (1957). Let Y_i be speaker i’s difference between the two condition means, divided by a declared standardizing scale. Thus positive values indicate longer closure in the first condition.
For this constructed example, treat the four speaker contrasts as independent and their response standard deviation as known to be one:
This expression is not the kernel of a familiar named family: the Student t prior is not conjugate to this normal likelihood, but that result changes the calculation rather than the definition of the posterior.
Z
=
\int_{-\infty}^{\infty}
\kappa(\theta';\mathbf{y})
\,\mathrm{d}\theta'.
The value Z is the normalizing constant. It does not vary with the particular \theta in the numerator, but it depends on the model and observed sample. Dividing by Z makes the posterior density integrate to one.
The unknown constant cancels. The Monte Carlo methods introduced on later pages can use relative posterior density, or its gradient, without first obtaining a formula for Z. The Stan documentation describes this computation as evaluation of a target log density up to an additive constant.
Numerical integration works here because the target is one dimensional. In a model with many parameters, direct numerical integration may require an impractical number of evaluations.
Evaluating on the log scale
Likelihood products can become numerically tiny. Computation thus usually evaluates
\log\kappa(\theta;\mathbf{y})
and forms density ratios by subtracting log densities. This representation changes neither the posterior target nor Z.
Separating computational difficulty from uncertainty
An unavailable closed-form expression for Z is a computational difficulty, but it does not make the posterior undefined when 0<Z<\infty. Under the model, f_{\Theta\mid\mathbf Y}(\theta\mid\mathbf{y}) remains the target.
The normalizing constant Z is a fixed mathematical consequence of the observed data and model. Numerical or Monte Carlo methods approximate posterior summaries and add computational error, which remains distinct from posterior spread and model adequacy.
A small Monte Carlo error still does not narrow the posterior or establish that the model fits the data.
Check your understanding
Which factor in the unnormalized posterior comes from the likelihood?
Why does Z cancel from a posterior density ratio?
Why is a missing closed-form normalizer not a missing posterior?
Why is the log scale useful for computation?
The computation sequence begins with Monte Carlo integration, where direct target draws are available.
---title: "When the posterior will not normalize by hand"---The [conjugate prior page](conjugate-priors.qmd) identified a normalized posterior by matching its algebraic form to a known distribution, but many useful prior and likelihood combinations do not supply that match. What, then, remains calculable?Let the nonnegative product of likelihood and prior have normalizing integral $Z$. A [**proper posterior**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html) is one that can be normalized to integrate to one, which occurs when $0<Z<\infty$. The product need not be positive at every parameter value, and a proper posterior can exist even when $Z$ has no closed-form expression.## Constructing a nonconjugate modelSuppose each of four speakers produces stops in two matched conditions. [Closure duration](https://www.cambridge.org/core/journals/language/article/abs/closure-duration-and-the-intervocalic-voicedvoiceless-distinction-in-english/397A7F377475B5423A000782430B755B) is the interval from the formation of a stop closure to its release, a measure used in phonetic work on distinctions among stops since at least Lisker (1957). Let $Y_i$ be speaker $i$'s difference between the two condition means, divided by a declared standardizing scale. Thus positive values indicate longer closure in the first condition.For this constructed example, treat the four speaker contrasts as independent and their response standard deviation as known to be one:$$Y_i\mid\Theta=\theta\sim\operatorname{Normal}(\theta,1).$$Let the sample mean be $\overline{y}=.50$. Terms not involving $\theta$ can be removed from the normal likelihood, leaving$$f_{\mathbf Y\mid\Theta}(\mathbf{y}\mid\theta)\propto\exp\left[-\frac{4}{2}(\theta-.50)^2\right].$$Choose a Student t prior with three degrees of freedom, location zero, and scale one:$$\Theta\sim t_3(0,1).$$Its density [**kernel**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html) is the part that varies with $\theta$ after constant factors have been omitted:$$f_\Theta(\theta)\propto\left(1+\frac{\theta^2}{3}\right)^{-2}.$$Multiplication gives the posterior kernel $\kappa$, a nonnegative function proportional to the posterior density:$$\kappa(\theta;\mathbf{y})=\exp\left[-2(\theta-.50)^2\right]\left(1+\frac{\theta^2}{3}\right)^{-2}.$$This expression is not the kernel of a familiar named family: the Student t prior is not conjugate to this normal likelihood, but that result changes the calculation rather than the definition of the posterior.## Identifying the missing quantityThe normalized posterior is$$f_{\Theta\mid\mathbf Y}(\theta\mid\mathbf{y})=\frac{\kappa(\theta;\mathbf{y})}{Z},$$where$$Z=\int_{-\infty}^{\infty}\kappa(\theta';\mathbf{y})\,\mathrm{d}\theta'.$$The value $Z$ is the [**normalizing constant**](https://bruno.nicenboim.me/bayescogsci/ch-introBDA.html). It does not vary with the particular $\theta$ in the numerator, but it depends on the model and observed sample. Dividing by $Z$ makes the posterior density integrate to one.## Comparing parameter values without $Z$For two values $\theta_a$ and $\theta_b$,$$\frac{f_{\Theta\mid\mathbf Y}(\theta_b\mid\mathbf{y})}{f_{\Theta\mid\mathbf Y}(\theta_a\mid\mathbf{y})}=\frac{\kappa(\theta_b;\mathbf{y})}{\kappa(\theta_a;\mathbf{y})}.$$The unknown constant cancels. The Monte Carlo methods introduced on later pages can use relative posterior density, or its gradient, without first obtaining a formula for $Z$. The [Stan documentation](https://mc-stan.org/docs/stan-users-guide/how-stan-works.html#target-log-density) describes this computation as evaluation of a target log density up to an additive constant.```{r}#| label: nonconjugate-duration-posterior#| echo: truelog_unnormalized <-function(theta) {-2* (theta - .50)^2+dt(theta, df =3, log =TRUE)}relative_density <-exp(log_unnormalized(.50) -log_unnormalized(0))normalizer <-integrate(function(theta) exp(log_unnormalized(theta)),lower =-Inf,upper =Inf)$valuestopifnot(relative_density >1)stopifnot(normalizer >0)c(relative_density_at_half_vs_zero = relative_density,normalizer = normalizer)```Numerical integration works here because the target is one dimensional. In a model with many parameters, direct numerical integration may require an impractical number of evaluations.## Evaluating on the log scaleLikelihood products can become numerically tiny. Computation thus usually evaluates$$\log\kappa(\theta;\mathbf{y})$$and forms density ratios by subtracting log densities. This representation changes neither the posterior target nor $Z$.## Separating computational difficulty from uncertaintyAn unavailable closed-form expression for $Z$ is a computational difficulty, but it does not make the posterior undefined when $0<Z<\infty$. Under the model, $f_{\Theta\mid\mathbf Y}(\theta\mid\mathbf{y})$ remains the target.The normalizing constant $Z$ is a fixed mathematical consequence of the observed data and model. Numerical or Monte Carlo methods approximate posterior summaries and add computational error, which remains distinct from posterior spread and model adequacy.A small Monte Carlo error still does not narrow the posterior or establish that the model fits the data.## Check your understanding1. Which factor in the unnormalized posterior comes from the likelihood?2. Why does $Z$ cancel from a posterior density ratio?3. Why is a missing closed-form normalizer not a missing posterior?4. Why is the log scale useful for computation?The computation sequence begins with [Monte Carlo integration](monte-carlo-integration.qmd), where direct target draws are available.