From outcomes to distributions
The foundations module gave us a probability space \langle\Omega,\mathcal F,\mathbb{P}\rangle. But a linguistic outcome in \Omega may be a complete token, utterance, or annotated tree, while the next analysis may need only its duration, case, or dependency length. How do we move from the complete outcome to the particular property named in the research question?
A random variable supplies this outcome-to-value mapping. Basically, it extracts a value without replacing the outcome that supports it. More specifically, it is a measurable function from outcomes to values, so questions about those values correspond to events in \mathcal F.
For a four-form pronoun space, for instance, a case variable can be defined by
C(\omega) \equiv \begin{cases} \mathrm{nom},&\omega\in\{\textit{he},\textit{they}\},\\ \mathrm{acc},&\omega\in\{\textit{him},\textit{them}\}. \end{cases}
Its value space and sigma-algebra are
\mathcal X_C\equiv\{\mathrm{nom},\mathrm{acc}\}, \qquad \mathcal G_C\equiv2^{\mathcal X_C}.
Now take any set of case values B\in\mathcal G_C. The event we can assign probability to is the set of pronoun outcomes that C sends into B, that is, the preimage C^{-1}(B). Thus the distribution of C assigns B the probability
\mathbb{P}(C\in B) \equiv \mathbb{P}\!\left(C^{-1}(B)\right).
As a function of B, this assignment is a probability measure on \langle\mathcal X_C,\mathcal G_C\rangle. We retain \mathbb{P} for the probability measure on \langle\Omega,\mathcal F\rangle and introduce no second measure symbol. Expressions such as \mathbb{P}(C=c) abbreviate the corresponding preimage events. Lowercase p, q, and r denote PMFs; lowercase f, g, and h denote PDFs; uppercase F and G denote CDFs. These conventions stay fixed throughout the module.
Random variables and one-variable distributions
We first explain why random variables are needed and then give the measurable-function definition. These pages assume the foundations treatment of event spaces and sigma-algebras. The next pages introduce discrete random variables and their probability mass functions, followed by absolutely continuous random variables and their probability density functions. A page on cumulative distribution functions then gives the representation shared by both types. This order states what values are possible and immediately specifies how probability is assigned to them.
Expectations and moments
Once a distribution has been represented, we can ask what features of it one number might summarize. The expected value is an integral with respect to a probability distribution when that integral exists. The notes then develop expectations of functions, linearity of expectation, and failure of finite expectations. Variance, standard deviation, and central moments summarize specified powers of centered values. Each summary assumes the relevant expectation exists.
Distribution families
We next reuse those objects in particular observation models. The discrete families are categorical, Bernoulli, binomial, hypergeometric, geometric, negative binomial, and Poisson. Their supports and probability mass functions differ because they represent different observation processes. The page order moves from one finite-category draw to totals, waits, and exposure-based counts.
The continuous families are uniform, beta, normal, chi-squared, and Student’s t. Each page states the support, parameterization, and density or defining construction before using the family. The sampling pages then show how a CDF turns a basic uniform draw into a draw from another distribution and how a random seed makes the computational realization reproducible.
Joint distributions
Finally, we move from one variable to several. A joint distribution preserves value pairings from the same outcome. Marginalization sums or integrates out one coordinate. Conditional independence states that one conditional distribution is invariant to an additional conditioning variable. Conditional expectation, covariance, and correlation summarize features of conditional or joint distributions. These pages rely on the foundations definitions of joint event probability, conditional probability, and independence, so they give only working reminders of those ideas.
The chapter dependency graph records the prerequisite relations among these definitions. Begin with why random variables are needed.