Code
pronoun_mass <- c(he = .20, him = .30, they = .10, them = .40)
all(pronoun_mass >= 0)
sum(pronoun_mass)
sum(pronoun_mass[c("him", "them")])The sample space \Omega specifies which outcomes the model represents, and the sigma-algebra \mathcal{F} specifies which events are measurable. We now have the objects we want to measure. How do we assign numbers to them without producing contradictions?
Basically, a probability measure assigns a number between zero and one to every measurable event. More specifically, a probability measure is a function
\mathbb{P}:\mathcal{F}\rightarrow[0,1]
that satisfies the three probability axioms stated below. Its arguments are events in \mathcal F, not unstructured labels or raw numerical values.
The sample space, sigma-algebra, and probability measure must be specified together. The triple
\langle\Omega,\mathcal{F},\mathbb{P}\rangle
is a probability space.
The fourteen-outcome example on the generating-event-spaces page uses the event space \sigma(T_{14},A_{14}). For every E\in\sigma(T_{14},A_{14}), define
\mathbb P(E) \equiv \frac{|E|}{|\Omega_{14}|} =\frac{|E|}{14}.
The formula gives every outcome equal weight when it counts the members of a measurable event. Note, though, that \sigma(T_{14},A_{14}) need not contain every singleton event. We may use singleton weights to compute \mathbb P(E) without claiming that every singleton is itself in the domain \mathcal F.
For instance, the third-person accusative event is
T_{14}\cap A_{14} =\{\textit{them},\textit{it}_{[+\mathrm{acc}]}, \textit{her},\textit{him}\}.
Thus
\mathbb P(T_{14}\cap A_{14}) =\frac{4}{14} =\frac{2}{7}.
The uniform assignment is one probability measure on the measurable space. It is not forced by the sample space or sigma-algebra. This distinction matters: \mathcal F says which questions can receive probabilities, while \mathbb P supplies their numerical answers.
Return to the four pronoun forms
\Omega_4 \equiv\{\textit{he},\textit{him},\textit{they},\textit{them}\}.
For this finite example, define \mathcal{F}\equiv2^{\Omega_4} so that every subset is measurable. Suppose a toy corpus motivates these assignments.
| outcome | assigned probability |
|---|---|
| he | .20 |
| him | .30 |
| they | .10 |
| them | .40 |
The table assigns probabilities to singleton events. The probability of a larger event is the sum of the probabilities assigned to its constituent outcomes.
Let
A_4\equiv\{\textit{him},\textit{them}\}
be the accusative event. Its probability is
\begin{aligned} \mathbb{P}(A_4) &=\mathbb{P}(\{\textit{him}\}) +\mathbb{P}(\{\textit{them}\})\\ &=.30+.40\\ &=.70. \end{aligned}
This probability measure assigns .70 to A_4. A different measure on the same measurable space may assign a different value.
Not every numerical assignment is a probability measure. Three requirements solve what we might call the coherence problem: they keep event probabilities nonnegative, normalized, and additive across disjoint alternatives.
First, every event has nonnegative probability:
\mathbb{P}(B)\geq 0 \qquad\text{for every }B\in\mathcal{F}.
Second, the full sample space has probability one:
\mathbb{P}(\Omega)=1.
Third, probabilities add across a countable collection of events that share no outcomes. If B_1,B_2,\ldots are pairwise disjoint, then
\mathbb{P}\!\left(\bigcup_i B_i\right) =\sum_i\mathbb{P}(B_i).
Each value in the table is nonnegative, so the first requirement holds for the singleton events. The singleton events divide \Omega_4 into nonoverlapping parts. Their values sum to
.20+.30+.10+.40=1,
so the probability of \Omega_4 is one.
Now combine he and him. Since the two singleton events share no outcomes,
\mathbb{P}(\{\textit{he},\textit{him}\}) =.20+.30 =.50.
By countable additivity, the singleton probabilities determine the probability of the larger event. We do not need another estimate for \{\textit{he},\textit{him}\}.
A few base R checks reproduce the arithmetic.
pronoun_mass <- c(he = .20, him = .30, they = .10, them = .40)
all(pronoun_mass >= 0)
sum(pronoun_mass)
sum(pronoun_mass[c("him", "them")])The three results are TRUE, 1, and .7.
The impossible event has probability zero. This is not a fourth assumption; it follows from the three requirements. Since \Omega and \varnothing are disjoint and \Omega\cup\varnothing=\Omega,
\mathbb{P}(\Omega) =\mathbb{P}(\Omega)+\mathbb{P}(\varnothing).
Subtracting \mathbb{P}(\Omega) from both sides gives
\mathbb{P}(\varnothing)=0.
The probability of a complement also follows from the requirements rather than being assumed separately. Since B and B^c are disjoint and their union is \Omega,
\mathbb{P}(B)+\mathbb{P}(B^c)=1.
Thus
\mathbb{P}(B^c)=1-\mathbb{P}(B).
For the accusative event, this gives \mathbb{P}(A_4^c)=1-.70=.30, which agrees with the mass assigned to he and they.
An assignment that violates any of the three requirements is not a probability measure. Suppose the four pronoun values were .20, .30, .10, and .50. Though each value lies between zero and one, their sum is 1.10, so the assignment violates normalization.
Renormalizing the values might produce a probability measure, but it changes every numerical claim. Before doing so, we should ask why the original values did not sum to one. They may be raw corpus counts, percentages computed from different denominators, or model scores that were never intended as probabilities.
The exact upshot is that a probability measure is a coherent numerical assignment to measurable events. The next page uses additivity to ask when two event probabilities can be added without counting any outcome twice.