Code
depth <- c(omega1 = 1, omega2 = 2, omega3 = 2)
names(depth)[depth == 2]
names(depth)[depth <= 2]The preceding page introduced a random variable as a function from complete outcomes to values. But not every function between arbitrary measurable spaces automatically supports probability statements. What property must the mapping have? We now add the measurability condition that answers this question.
Suppose the outcomes in a syntax study are fully annotated sentence trees. The analysis concerns embedding depth, so we define
H(\omega)\equiv\text{maximum embedding depth in tree }\omega.
The function takes a complex outcome and returns a value. To use that value in probability statements, every measurable set of values must lead back to an event that the original model can measure.
This measurability requirement completes the formal definition of the random variable introduced on the preceding page.
Let \langle\Omega,\mathcal{F}\rangle be the measurable space of sentence tree outcomes, using the event space introduced earlier. Let \langle\mathcal X,\mathcal{G}\rangle be the measurable space of depth values. The function has the form
H:\Omega\longrightarrow\mathcal X.
The set \mathcal X is the declared codomain of H, written \operatorname{cod}(H). A statement such as x\in\operatorname{cod}(X) refers to the declared value space. It does not mean that x appeared in the observed sample, and it does not require that x have positive probability.
For a small finite study, the value set might be
\mathcal X\equiv\{0,1,2,3\}.
The sigma-algebra \mathcal{G} contains the measurable sets of depth values. If every subset is measurable, then \{2\} and \{2,3\} both belong to \mathcal{G}.
For this illustration, let \Omega\equiv\{\omega_1,\omega_2,\omega_3\}. The outcomes receive these values.
| sentence tree outcome | H(\omega) |
|---|---|
| \omega_1 | 1 |
| \omega_2 | 2 |
| \omega_3 | 2 |
Select the value set \{2\}. Before writing notation, we can state the operation in words: collect every sentence tree whose depth value is 2. The outcomes mapped into that set are
H^{-1}(\{2\}) \equiv\{\omega\in\Omega\mid H(\omega)\in\{2\}\} =\{\omega_2,\omega_3\}.
This set is the preimage of \{2\} under H.
The superscript -1 does not say that H has an inverse function. Both \omega_2 and \omega_3 map to 2, so the depth value does not recover one unique tree. The preimage collects every outcome whose mapped value belongs to the selected set.
For the set \{1,2\},
H^{-1}(\{1,2\})=\{\omega_1,\omega_2,\omega_3\}.
The operation always moves from a set of values back to a set of outcomes. This backward move is what lets the original probability measure answer a question stated in the value space.
The function H is a random variable when
E\in\mathcal{G} \quad\Longrightarrow\quad H^{-1}(E)\in\mathcal{F}.
This condition is called measurability. It says that every measurable question about the value corresponds to a measurable event in the original outcome space.
If \{2\}\in\mathcal{G}, then the event \{\omega_2,\omega_3\} must belong to \mathcal{F}. Otherwise the expression “the probability that embedding depth is 2” would refer to an event to which the model assigns no probability.
In a finite model with \mathcal{F}\equiv2^\Omega, every preimage is measurable because every subset of outcomes belongs to \mathcal{F}. The requirement becomes more substantive for larger outcome spaces, but its interpretation remains the same.
Use the case-coded pronoun outcomes
\begin{aligned} \Omega\equiv\{&\textit{I},\textit{me}, \textit{you}_{[+\mathrm{acc}]},\textit{you}_{[-\mathrm{acc}]}, \textit{they},\textit{them}, \textit{it}_{[+\mathrm{acc}]},\textit{it}_{[-\mathrm{acc}]},\\ &\textit{she},\textit{her},\textit{he},\textit{him}, \textit{we},\textit{us}\}. \end{aligned}
Define V:\Omega\to\{1,\ldots,14\} by
\begin{aligned} V(\textit{I})&\equiv1, &V(\textit{me})&\equiv2, &V(\textit{you}_{[+\mathrm{acc}]})&\equiv3, &V(\textit{you}_{[-\mathrm{acc}]})&\equiv4,\\ V(\textit{they})&\equiv5, &V(\textit{them})&\equiv6, &V(\textit{it}_{[+\mathrm{acc}]})&\equiv7, &V(\textit{it}_{[-\mathrm{acc}]})&\equiv8,\\ V(\textit{she})&\equiv9, &V(\textit{her})&\equiv10, &V(\textit{he})&\equiv11, &V(\textit{him})&\equiv12,\\ V(\textit{we})&\equiv13, &V(\textit{us})&\equiv14. \end{aligned}
For instance,
V^{-1}(\{1,2,3\}) =\{\textit{I},\textit{me},\textit{you}_{[+\mathrm{acc}]}\}
and
V^{-1}(\{12,13,14\}) =\{\textit{him},\textit{we},\textit{us}\}.
The enumeration is arbitrary. Its purpose is to demonstrate that a random variable maps individual outcomes to values, while a preimage maps a set of values back to an event.
Once H is defined, define probability notation for its values by
\mathbb{P}(H=2) \equiv\mathbb{P}\bigl(H^{-1}(\{2\})\bigr) =\mathbb{P}(\{\omega_2,\omega_3\}).
Similarly,
\mathbb{P}(H\leq2)
means the probability of the preimage of all depth values at or below 2. The notation is compact, but the probability measure still acts on an event of outcomes.
depth <- c(omega1 = 1, omega2 = 2, omega3 = 2)
names(depth)[depth == 2]
names(depth)[depth <= 2]The first result encodes H^{-1}(\{2\}). The second encodes the preimage of the set of values at or below 2.
One outcome cannot be recovered from a many to one mapping. The value 2 does not determine whether \omega_2 or \omega_3 occurred. The notation H^{-1}(\{2\}) denotes the preimage containing both outcomes rather than selecting one.
A duration, rating, or word count rarely identifies the full linguistic record from which it was extracted. Many outcomes can map to the same random-variable value.
We now know what a random variable must preserve. The next page considers variables whose possible values can be listed.