The preceding page approximated an expectation using draws from the target distribution itself. Importance sampling instead uses draws from a proposal distribution. Consider a constructed production study in which a token is coded one when the constituent bearing contrastive focus is produced with a pitch accent. This coding treats contrastive focus as an information-structural category and pitch accent as one possible prosodic realization, not as equivalent notions. Suppose the posterior target for the token probability is

\Theta \sim \operatorname{Beta}(8,4).

Write f(\theta) for the corresponding beta density. We want the posterior probability that \Theta>.70, but suppose direct target draws are unavailable. The question is whether draws from another distribution can still recover a target expectation. We can draw \Theta^{(s)} from a proposal distribution Q with density q(\theta) and correct the mismatch between proposal and target with weights.

Rewriting the target expectation

For a function g, the target expectation is

\mathbb{E}_f[g(\Theta)] = \int g(\theta)f(\theta)\,\mathrm{d}\theta.

Multiply and divide by a proposal density q(\theta):

\begin{aligned} \mathbb{E}_p[g(\Theta)] &= \int g(\theta) \frac{f(\theta)}{q(\theta)} q(\theta) \,\mathrm{d}\theta\\ &= \mathbb{E}_q \left[ g(\Theta)w(\Theta) \right], \end{aligned}

where

w(\theta)=\frac{f(\theta)}{q(\theta)}

is the importance weight. A proposal draw receives a large weight when it falls in a region that is relatively common under the target.

Using a uniform proposal

Choose the proposal distribution

Q=\operatorname{Uniform}(0,1), \qquad q(\theta)=\mathbf{1}\{0<\theta<1\}.

Since its density is one on the unit interval, the raw weights are beta densities evaluated at the uniform draws. When the target density is normalized, an ordinary importance sampling estimate is

\widehat{I}_S =\frac{1}{S}\sum_{s=1}^S w^{(s)}g(\theta^{(s)}).

We will use the self-normalized form, which divides the weighted sum by the sum of the weights. Define normalized weights

\widetilde{w}^{(s)} = \frac{w^{(s)}}{\sum_{r=1}^Sw^{(r)}}.

The self-normalized estimate is

\widehat{I}_S = \sum_{s=1}^S \widetilde{w}^{(s)}g(\theta^{(s)}).

For the tail probability, g(\theta)=\mathbf{1}\{\theta>.70\}.

Self-normalization generally introduces bias when S is finite. Under the proposal-support condition described below and the required integrability conditions, however, the estimate approaches the target expectation as S grows. This limiting property is called consistency. The method also works when the target is known only up to a constant. If

f(\theta)=\frac{\kappa(\theta)}{Z},

form raw weights as

w^{(s)} = \frac{\kappa(\theta^{(s)})} {q(\theta^{(s)})}.

The shared factor 1/Z disappears when these weights are divided by their sum. The calculation can thus approximate a normalized target expectation without evaluating Z.

Code
set.seed(964)
draw_count <- 50000
theta_proposal <- runif(draw_count)
raw_weight <- dbeta(theta_proposal, 8, 4) /
  dunif(theta_proposal)
weight <- raw_weight / sum(raw_weight)

importance_estimate <- sum(weight * (theta_proposal > .70))
exact_value <- 1 - pbeta(.70, 8, 4)

stopifnot(abs(sum(weight) - 1) < 1e-12)
stopifnot(abs(importance_estimate - exact_value) < .01)

c(importance = importance_estimate, exact = exact_value)

The normalized weights change an unweighted average under the uniform proposal into a weighted average approximating the beta target.

Inspecting weight concentration

Coverage is necessary. More precisely, the target must assign zero probability to every region that the proposal assigns zero probability. If f(\theta)>0 on a region of positive target probability where q(\theta)=0, no proposal draw can represent that region, and no weight can repair the omission.

Shared support is not sufficient for an efficient calculation. A proposal that puts very little probability in an important target region may produce a few enormous weights and many negligible weights. The estimate then depends on a small number of proposal draws.

One direct summary is the largest normalized weight:

Code
largest_weight <- max(weight)
weight_share_top_100 <- sum(sort(weight, decreasing = TRUE)[1:100])

stopifnot(largest_weight < .01)
c(largest_weight = largest_weight,
  share_in_largest_100 = weight_share_top_100)

These values do not prove that the estimate is accurate, but they reveal whether a very small part of the simulation dominates it.

Avoiding the proposal-frequency reading

The proposal distribution is a computational device, not a representation of prior beliefs, posterior uncertainty, or a linguistic data-generating process. Frequent proposal draws do not by themselves indicate posterior support. Only the weighted draws approximate the target.

Monte Carlo error also remains. More proposal draws may stabilize the estimate, while a better-matched proposal may improve it more efficiently. Neither change narrows the posterior target or repairs a poor linguistic model.

A realized set of moderate weights does not prove that every important target tail has been represented. Proposal choice should be justified from support and tail behavior, not only from one weight histogram.

Check your understanding

  1. Why does the ratio f(\theta)/q(\theta) correct proposal draws?
  2. What fails if the proposal has zero density in a target region?
  3. Why can a proposal with full support still be inefficient?
  4. Which distribution determines the unweighted draw frequency, and which determines the weighted target estimate?