Code
pi_overt <- .60
sequence_probability <- pi_overt^3 * (1 - pi_overt)
three_of_four_probability <- choose(4, 3) * sequence_probability
c(sequence_probability, three_of_four_probability)The preceding page distinguished the target population from the realized sample. We now need a probability model that connects repeated samples to a population parameter. Consider a constructed coding example in which we sample four clauses and record whether each has an overt subject. Here, overt means that the clause contains an expressed subject phrase rather than an unexpressed subject.
X_i= \begin{cases} 1 & \text{if clause }i\text{ has an overt subject},\\ 0 & \text{otherwise}. \end{cases}
To connect these observations to one population probability, we need a sampling model. A common starting point treats the observations as independent and identically distributed (IID). We will treat the two adjectives separately because they license different later steps.
Assume every clause has the same Bernoulli distribution:
X_i\sim\operatorname{Bernoulli}(\pi)
for every i. This says
\mathbb{P}(X_i=1)=\pi
for each observation. The overt subject probability does not change across clause positions under the model.
This is the identically distributed part of the IID assumption. It is a claim about the modeled PMF, not a claim that all observed clauses have the same realized value.
Now fix \pi and assume that the complete collection of clause outcomes is mutually independent. This is the independence condition introduced in the foundations module, applied to the sampling distribution at that parameter value:
p(x_1,\ldots,x_n\mid\pi) =\prod_{i=1}^n p(x_i\mid\pi).
Mutual independence is stronger than checking each pair separately. It is the assumption that justifies the product for the complete sample.
Together, the assumptions are written
X_1,\ldots,X_n \mathrel{\overset{\mathrm{IID}}{\sim}} \operatorname{Bernoulli}(\pi).
The parameter \pi is shared; the realized outcomes may differ.
Suppose the observed sequence is
(x_1,x_2,x_3,x_4)=(1,0,1,1).
The probability chain rule first gives
\begin{aligned} p(x_1,x_2,x_3,x_4\mid\pi) ={}&p(x_1\mid\pi)\\ &\times p(x_2\mid x_1,\pi)\\ &\times p(x_3\mid x_1,x_2,\pi)\\ &\times p(x_4\mid x_1,x_2,x_3,\pi). \end{aligned}
Independence removes the earlier observations from each later factor:
p(x_1,x_2,x_3,x_4\mid\pi) =\prod_{i=1}^{4}p(x_i\mid\pi).
For the observed sequence,
\begin{aligned} p(1,0,1,1\mid\pi) &=\pi(1-\pi)\pi\pi\\ &=\pi^3(1-\pi). \end{aligned}
The product is a consequence of the independence assumption. It is not an algebraic convenience that can be used regardless of sampling structure.
If \pi=.60, then
p(1,0,1,1\mid.60) =(.60)^3(.40) =.0864.
This is the probability of that exact ordered sequence. The probability of any sequence containing three overt subjects among four clauses is larger because there are four placements for the single zero.
pi_overt <- .60
sequence_probability <- pi_overt^3 * (1 - pi_overt)
three_of_four_probability <- choose(4, 3) * sequence_probability
c(sequence_probability, three_of_four_probability)The distinction between one sequence and an unordered count will matter when likelihoods are written from different data representations.
Clauses from one speaker may share that speaker’s subject-expression tendency, while clauses in one document may share genre, topic, or discourse structure. These repetitions can violate independence.
The identical distribution claim may also fail. Questions, subordinate clauses, and imperatives can have different overt subject probabilities even when sampled independently.
Ten clauses from one speaker are ten clause observations. They are not equivalent to ten independently sampled speakers when the population claim concerns speaker variation.
A unique key establishes that two rows are different observations. It does not make the observations independent or remove their shared speaker, item, or document.
Specify the proposed independent unit and list the identifiers repeated across rows. Later models represent those repeated sources directly.
The next page applies this sampling statement to a teaching extract while keeping its linguistic limitations explicit.