Code
syllable_value <- c("1", "2", "3", "4")
syllable_mass <- c(.42, .33, .17, .08)
all(syllable_mass >= 0)
sum(syllable_mass)
sum(syllable_mass[syllable_value %in% c("3", "4")])The preceding page distinguished a random variable’s declared codomain from its positive-probability support. Suppose S records the number of syllables in a sampled word token. Knowing that S is discrete does not tell us how likely its codomain values are. What object pairs each possible count with its probability?
Consider this constructed probability table.
The syllable counts and masses are hypothetical. In observed data, the count would depend on a declared pronunciation and syllabification analysis; the PMF cannot resolve that linguistic representation for us.
| syllable count s | probability |
|---|---|
| 1 | .42 |
| 2 | .33 |
| 3 | .17 |
| 4 | .08 |
Basically, we need a lookup rule from values to probabilities. More specifically, the function assigning a probability to each codomain value is a probability mass function, abbreviated PMF. We write
p_S:\mathbb Z_{\geq0}\to[0,1], \qquad p_S(s)\equiv\mathbb{P}(S=s).
The capital S denotes the random variable. The lowercase s stands for one possible value. Thus
p_S(2)=.33.
By the PMF definition, the probability measure assigns .33 to the event \{\omega\in\Omega\mid S(\omega)=2\}. In this constructed model, \operatorname{cod}(S)=\mathbb Z_{\geq0} and \operatorname{supp}(S)=\{1,2,3,4\}. Thus p_S(s)=0 for every other s\in\operatorname{cod}(S).
First, a PMF must assign nonnegative mass:
p_S(s)\geq0 \qquad\text{for every }s\in\operatorname{cod}(S).
Second, the masses across the codomain must sum to one:
\sum_{s\in\operatorname{cod}(S)}p_S(s)=1.
Values in the codomain but outside the support contribute zero, so this sum has the same numerical value as a sum restricted to \operatorname{supp}(S) without using support as the default indexing set.
For the constructed table,
.42+.33+.17+.08=1.
The rows represent the positive-probability value events. They are mutually exclusive, and their union has probability one. The other codomain values remain part of the PMF but contribute zero mass.
What is the probability of a token with at least three syllables? The relevant event contains the rows for 3 and 4:
\begin{aligned} \mathbb{P}(S\geq3) &=p_S(3)+p_S(4)\\ &=.17+.08\\ &=.25. \end{aligned}
The calculation adds masses because the two value events do not overlap. A token cannot have exactly three syllables and at least four syllables under this representation.
syllable_value <- c("1", "2", "3", "4")
syllable_mass <- c(.42, .33, .17, .08)
all(syllable_mass >= 0)
sum(syllable_mass)
sum(syllable_mass[syllable_value %in% c("3", "4")])The three results are TRUE, 1, and .25. These checks establish that the entered table satisfies the PMF requirements and recover the selected event probability.
If p_S(7)=0, then the event S=7 has zero probability under the stated model. This may reflect a deliberately restricted support or a model that treats the value as impossible. It need not mean that no seven syllable word can exist in every linguistic population.
The PMF belongs to a represented sampling process. A lexicon weighted by word types and a corpus weighted by word tokens can assign different masses to the same syllable counts.
A probability mass function cannot omit codomain values without accounting for their probability. Suppose the first three positive-mass rows above were retained but the final row were dropped. The remaining masses would sum to .92, so the table would not define a PMF unless the omitted value were assigned its missing mass elsewhere.
Renormalizing the three values would define a different model conditional on at most three syllables. That change may be useful, but it must be stated as a new reference event rather than treated as harmless deletion.
A PMF assigns mass to separate support values. The next page considers a model whose values instead fill an interval.