Code
joint_tense_number <- matrix(
c(.30, .15, .20, .35),
nrow = 2,
dimnames = list(
tense = c("past", "present"),
number = c("singular", "plural")
)
)
rowSums(joint_tense_number)
colSums(joint_tense_number)The preceding page defined a distribution over paired values. Suppose a morphological analysis records tense T and number N for one sampled verb token. If our next question concerns tense alone, how do we remove the number distinction without discarding any token probability?
The four cell values are constructed annotations. A real application must state how tense and number are analyzed, including how syncretic or underspecified forms are handled; marginalization cannot resolve those annotation choices.
| singular N=\mathrm{sg} | plural N=\mathrm{pl} | |
|---|---|---|
| past T=\mathrm{past} | .30 | .20 |
| present T=\mathrm{pres} | .15 | .35 |
Basically, we hold one tense row fixed and add across all number columns. No probability disappears; only the number distinction does. More specifically, for a fixed tense value t, the events \{\omega\in\Omega\mid T(\omega)=t,\ N(\omega)=n\} are disjoint as n varies, and
\{\omega\in\Omega\mid T(\omega)=t\} =\biguplus_{n\in\operatorname{cod}(N)} \{\omega\in\Omega\mid T(\omega)=t,\ N(\omega)=n\}.
The union is disjoint because one token cannot have both number values under the stated annotation. The probability-addition rule thus gives
\mathbb{P}(T=t) =\sum_{n\in\operatorname{cod}(N)} \mathbb{P}(T=t,N=n).
The PMFs are defined by
p_T(t)\equiv\mathbb{P}(T=t)
and
p_{T,N}(t,n)\equiv\mathbb{P}(T=t,N=n).
Substitution yields the marginalization identity
p_T(t) =\sum_{n\in\operatorname{cod}(N)}p_{T,N}(t,n).
This derived one-variable PMF specifies the marginal distribution of T.
The past tense mass is
\begin{aligned} p_T(\mathrm{past}) &=p_{T,N}(\mathrm{past},\mathrm{sg}) +p_{T,N}(\mathrm{past},\mathrm{pl})\\ &=.30+.20\\ &=.50. \end{aligned}
The present tense mass is
p_T(\mathrm{pres})=.15+.35=.50.
Thus the marginal tense PMF is (.50,.50).
The same operation down columns gives the marginal number PMF:
p_N(\mathrm{sg})=.30+.15=.45
and
p_N(\mathrm{pl})=.20+.35=.55.
| singular | plural | p_T(t) | |
|---|---|---|---|
| past | .30 | .20 | .50 |
| present | .15 | .35 | .50 |
| p_N(n) | .45 | .55 | 1 |
For variables with joint density f_{X,Y}, define the marginal density of X by
f_X(x) \equiv\int_{-\infty}^{\infty}f_{X,Y}(x,y)\,\mathrm{d}y.
The integral ranges over the complete Y support while x remains fixed. For every measurable set A,
\begin{aligned} \mathbb{P}(X\in A) &=\int_A f_X(x)\,\mathrm{d}x\\ &=\int_A\!\left[ \int_{-\infty}^{\infty}f_{X,Y}(x,y)\,\mathrm{d}y \right]\mathrm{d}x. \end{aligned}
The inner integral sums out Y; the outer integral assigns probability to the event \{\omega\in\Omega\mid X(\omega)\in A\}.
joint_tense_number <- matrix(
c(.30, .15, .20, .35),
nrow = 2,
dimnames = list(
tense = c("past", "present"),
number = c("singular", "plural")
)
)
rowSums(joint_tense_number)
colSums(joint_tense_number)The row sums give the tense margin; the column sums give the number margin.
Marginalization sums over every value of one variable. Conditioning instead fixes one value and renormalizes, using the conditional-probability definition. For p_N(n)>0, define
p_{T\mid N}(t\mid n) \equiv \mathbb P(T=t\mid N=n) =\frac{p_{T,N}(t,n)}{p_N(n)}.
For plural tokens,
\begin{aligned} p_{T\mid N}(\mathrm{past}\mid\mathrm{pl}) &=\frac{p_{T,N}(\mathrm{past},\mathrm{pl})}{p_N(\mathrm{pl})}\\ &=\frac{.20}{.55}\\ &\approx.364. \end{aligned}
Rearranging the definition yields the probability factorization
p_{T,N}(t,n)=p_{T\mid N}(t\mid n)p_N(n).
Marginalizing number gives the tense distribution without conditioning on a particular number value, whereas conditioning on N=\mathrm{pl} gives a different distribution, here (.364,.636). Neither operation asserts independence.
In the joint table, plural tokens are more often present than past, while singular tokens are more often past than present. Those pairings disappear from the tense margin (.50,.50).
The marginal distribution is determined by the joint distribution, but the converse is false. The two one-variable marginals do not determine the cell probabilities or the association between T and N.
Marginalization removes a distinction. The next page asks whether two variables remain associated after a third distinction is held fixed.