The Fisher’s exact test page constructed a finite conditional reference distribution for a two by two table. The chi-squared test of independence instead uses a large-sample reference approximation. Consider a constructed corpus table that classifies 300 nominal-expression tokens by grammatical role and expression type. Here, lexical noun phrase means a noun-headed phrase rather than a personal pronoun; it does not mean an orthographic word form. The table assumes that nonreferential uses, such as expletive it, have been excluded. The form labels alone would not establish that exclusion in a real corpus.
Every cell contributes a nonnegative quantity. For the subject-pronoun cell,
\frac{(120-100)^2}{100}=4.
For the subject-lexical-noun-phrase cell,
\frac{(30-50)^2}{50}=8.
The object row contributes the same two values:
\frac{(80-100)^2}{100}=4
and
\frac{(70-50)^2}{50}=8.
Thus
\chi^2=4+8+4+8=24.
Larger values record larger aggregate discrepancies from the independence table.
Determining the reference distribution
For a table with r rows and c columns, the chi-squared reference distribution has
\nu=(r-1)(c-1)
degrees of freedom. Here,
\nu=(2-1)(2-1)=1.
Once the margins are fixed, only one cell can vary freely. Its value determines the other three counts, which gives the same one degree of freedom seen in the formula.
The p value is approximately 9.63\times10^{-7}. The observed allocation would be highly unusual under the approximate independence reference model.
Localizing the discrepancy
The total statistic is omnibus. It shows that the complete table departs from independence, but it does not say that every imaginable comparison differs.
The signed differences reveal the pattern: subjects have 20 more personal pronouns and 20 fewer lexical noun phrases than expected, while objects show the reverse. Because each deviation is squared in \chi^2, the total statistic itself discards direction.
An omnibus rejection does not identify which cells produce the association. Inspect observed counts, expected counts, and signed residuals to state where the association lies.
Checking the approximation and independence
The chi-squared reference distribution is a large-sample approximation. It tends to work better when expected cell counts are not sparse. Every expected count here is at least 50, so sparsity is not the main concern.
Observation independence is separate. Repeated expressions from the same speaker or document can be dependent even when all expected counts are large. Large counts do not repair clustered sampling.
For a sparse two-by-two table, Fisher’s exact test supplies a finite-sample conditional calculation. If multiple predictors, interactions, or repeated units matter, a regression model can represent more of the linguistic structure than either table test.
Check your understanding
Calculate the expected object-pronoun count from the margins.
Why does a two-by-two table have one degree of freedom once its margins are fixed?
Which cells contribute most to \chi^2, and why?
Why can document clustering invalidate the test even when every expected count is large?
The next section develops posterior distributions, which represent uncertainty about parameter values using a different inferential framework.
---title: "The chi-squared test of independence"---The [Fisher's exact test page](fishers-exact-test.qmd) constructed a finite conditional reference distribution for a two by two table. The chi-squared test of independence instead uses a large-sample reference approximation. Consider a constructed corpus table that classifies 300 nominal-expression tokens by grammatical role and expression type. Here, *lexical noun phrase* means a noun-headed phrase rather than a personal pronoun; it does not mean an orthographic word form. The table assumes that nonreferential uses, such as expletive *it*, have been excluded. The form labels alone would not establish that exclusion in a real corpus.|| personal pronoun | lexical noun phrase | total ||---|---:|---:|---:|| subject | 120 | 30 | 150 || object | 80 | 70 | 150 || total | 200 | 100 | 300 |The [**chi-squared test of independence**](https://openstax.org/books/introductory-business-statistics-2e/pages/11-4-test-of-independence) compares each observed cell count with the count expected if grammatical role and referring-expression type satisfied the [independence condition introduced earlier](../foundations/independence.qmd). The calculation thus begins by constructing the complete table implied by that condition.## Constructing one expected countUnder independence,$$\mathbb{P}(\text{subject and pronoun})=\mathbb{P}(\text{subject})\mathbb{P}(\text{pronoun}).$$The marginal probabilities estimated from the table are$$\widehat{\mathbb{P}}(\text{subject})=\frac{150}{300}=.5$$and$$\widehat{\mathbb{P}}(\text{pronoun})=\frac{200}{300}=\frac{2}{3}.$$Multiplying this joint probability by the grand total converts a probability for one sampled expression into the expected count among 300 expressions:$$E_{11}=300(.5)\left(\frac{2}{3}\right)=100.$$The same calculation can be written directly from the margins:$$E_{ij}=\frac{(\text{row }i\text{ total})\times(\text{column }j\text{ total})}{\text{grand total}}.$$The expected table under independence is|| personal pronoun | lexical noun phrase | total ||---|---:|---:|---:|| subject | 100 | 50 | 150 || object | 100 | 50 | 150 || total | 200 | 100 | 300 |The expected table has the same margins as the observed table. Only the allocation within rows differs.## Adding the cell discrepanciesPearson's statistic is$$\chi^2=\sum_i\sum_j\frac{(O_{ij}-E_{ij})^2}{E_{ij}}.$$Every cell contributes a nonnegative quantity. For the subject-pronoun cell,$$\frac{(120-100)^2}{100}=4.$$For the subject-lexical-noun-phrase cell,$$\frac{(30-50)^2}{50}=8.$$The object row contributes the same two values:$$\frac{(80-100)^2}{100}=4$$and$$\frac{(70-50)^2}{50}=8.$$Thus$$\chi^2=4+8+4+8=24.$$Larger values record larger aggregate discrepancies from the independence table.## Determining the reference distributionFor a table with $r$ rows and $c$ columns, the chi-squared reference distribution has$$\nu=(r-1)(c-1)$$degrees of freedom. Here,$$\nu=(2-1)(2-1)=1.$$Once the margins are fixed, only one cell can vary freely. Its value determines the other three counts, which gives the same one degree of freedom seen in the formula.The upper-tail probability is```{r}#| label: chi-squared-reference-table#| echo: truereference_table <-matrix(c(120, 30,80, 70),nrow =2,byrow =TRUE)result <-chisq.test(reference_table, correct =FALSE)expected <- result$expectedcontributions <- (reference_table - expected)^2/ expectedp_value <-pchisq(24, df =1, lower.tail =FALSE)stopifnot(isTRUE(all.equal(as.numeric(expected), c(100, 100, 50, 50))))stopifnot(isTRUE(all.equal(sum(contributions), 24)))stopifnot(abs(p_value -9.63357e-07) <0.000000000001)expectedcontributionsc(chisq =unname(result$statistic), p = p_value)```The p value is approximately $9.63\times10^{-7}$. The observed allocation would be highly unusual under the approximate independence reference model.## Localizing the discrepancyThe total statistic is omnibus. It shows that the complete table departs from independence, but it does not say that every imaginable comparison differs.The signed differences reveal the pattern: subjects have 20 more personal pronouns and 20 fewer lexical noun phrases than expected, while objects show the reverse. Because each deviation is squared in $\chi^2$, the total statistic itself discards direction.An omnibus rejection does not identify which cells produce the association. Inspect observed counts, expected counts, and signed residuals to state where the association lies.## Checking the approximation and independenceThe chi-squared reference distribution is a large-sample approximation. It tends to work better when expected cell counts are not sparse. Every expected count here is at least 50, so sparsity is not the main concern.Observation independence is separate. Repeated expressions from the same speaker or document can be dependent even when all expected counts are large. Large counts do not repair clustered sampling.For a sparse two-by-two table, [Fisher's exact test](fishers-exact-test.qmd) supplies a finite-sample conditional calculation. If multiple predictors, interactions, or repeated units matter, a regression model can represent more of the linguistic structure than either table test.## Check your understanding1. Calculate the expected object-pronoun count from the margins.2. Why does a two-by-two table have one degree of freedom once its margins are fixed?3. Which cells contribute most to $\chi^2$, and why?4. Why can document clustering invalidate the test even when every expected count is large?The next section develops [posterior distributions](posterior-distributions.qmd), which represent uncertainty about parameter values using a different inferential framework.