From linguistic records to probabilities
Suppose a speaker says heed once. How many observations do we have? The answer could be one recording, one word token, one vowel token, or several acoustic measurements taken during the vowel. Each answer produces a different dataset, even though every answer begins with the same utterance.
This choice creates the representation problem: before we count anything, we must decide what one row represents and which differences among speakers, items, and tokens remain visible. The same problem arises with every kind of linguistic data. A corpus can be divided into documents, sentences, word tokens, or word types. A reading-time experiment can retain one row for each trial or average those trials by participant or sentence.
This module works outward from that problem. We first decide what the observations are, then decide which comparisons answer a linguistic question, and only then introduce probability as a way of assigning numerical weight to possible outcomes and events.
Linguistic data and comparisons
First, we turn linguistic records into data. A ten-minute interview is one audio file, but it contains many word and vowel tokens. Once we distinguish the file from the observations extracted from it, we can state what each row represents, which columns identify it, and which speakers or words occur in more than one row.
Second, we turn a research question into a comparison. If an aspectually mismatched sentence takes 45 milliseconds longer to read than a matched sentence, we need to know which response was measured, how the conditions differed, and whether participants and sentence items occurred in both conditions. These facts determine which difference we have estimated.
Third, we work through Zipf’s law by counting word tokens, grouping them into types, ordering the types by frequency, and comparing the observed counts with an inverse rank function. A rank frequency function may describe a corpus well, but that fit does not tell us which process produced the corpus.
Probability
Now we can state the probability question. What results can the observation process produce, which groupings of those results matter, and how much probability should each grouping receive? A sample space is the set of possible outcomes represented by the model. An event is a set of those outcomes. If the outcomes are the forms he, him, they, and them, for instance, the plural event contains they and them.
We do not always assign probabilities to every subset of the sample space. A sigma-algebra specifies which events are measurable. We will construct the sigma-algebra generated by number and case for the four pronoun forms, then define a probability measure that assigns a probability to each measurable event. The sample space, sigma-algebra, and probability measure form the probability space
\langle\Omega,\mathcal{F},\mathbb{P}\rangle.
Once the probability space is defined, we can calculate probabilities for relations among events. We begin with mutually exclusive events, whose intersection is empty, and introduce the standard comma notation for joint probabilities. We then define conditional probability by restricting probability to a conditioning event. The chain rule follows by rearranging that definition, and Bayes’ rule follows by factoring the same joint probability in two orders. We conclude with independence: its conditional interpretation is invariance under conditioning, and its unrestricted definition is the product factorization.
The chapter dependency graph shows which later chapters use these definitions. The order matters: each page adds one object that the next page assumes. Begin by deciding what one linguistic observation represents.