Fisher’s exact test

The hypergeometric distribution page assigned probabilities when sampling without replacement from a finite population. Fisher’s exact test uses the same distribution after conditioning on the margins of a contingency table. A contingency table records joint counts for two categorical variables, and its margins are the row and column totals.

Consider a constructed sample of noun lexemes classified by the syllable count of the singular lemma and by plural-pattern class. A lexeme is an abstract vocabulary item. Here, we use lemma for the singular citation form used to represent it. The plural classes are a productive -(e)s realization and other realizations, such as stem change or suppletion, in which a different stem expresses the contrast. These classes do not refer to the phonetic variants of the regular suffix.

productive -(e)s plural other plural realization total
monosyllabic 18 12 30
polysyllabic 26 4 30
total 44 16 60

The statistical question is whether lemma syllable count and plural-pattern class are independent in the population represented by the sampling process. Fisher’s exact test constructs a conditional null distribution over possible tables while holding the observed row and column totals fixed. We first determine which tables remain possible under those margins, then assign each table a null probability, and finally decide which probabilities enter the tail calculation.

Finding the one free cell

Once all four margins are fixed, choosing the upper-left count, which we denote by A, determines the remaining cells. The observed value is A=18.

If A=20, the first row must still total 30, so its other-realization count is 10; the productive-pattern column must still total 44, so the lower-left count is 24; and the final cell must then be 6.

productive -(e)s plural other plural realization total
monosyllabic 20 10 30
polysyllabic 24 6 30
total 44 16 60

The smallest possible A is 14 because the lower-left cell cannot exceed its row total of 30. The largest possible A is 30 because the first row contains only 30 nouns. Thus the reference set under the null hypothesis contains tables with A=14,15,\ldots,30.

Assigning null probabilities

Under independence with fixed margins, A follows a hypergeometric distribution. There are {60\choose30} ways to select the 30 noun lexemes assigned to the monosyllabic row. The observed table places 18 of the 44 productive-pattern lexemes and 12 of the 16 other-pattern lexemes in that row. Its probability is

\mathbb{P}_{H_0}(A=18) = \frac{{44\choose18}{16\choose12}} {{60\choose30}} \approx.01584.

This is the probability of the observed table under the conditional null model. It is not yet a p value because the p value must also include other tables judged at least as incompatible with the null.

Forming the two-sided p value

A one-sided alternative supplies a direction and lets us sum one tail. A two-sided alternative needs a rule for identifying tables at least as discrepant as the observed table.

R’s fisher.test() includes every table whose null probability is no greater than the probability of the observed table. This is a probability-ordering definition of the two-sided p value.

Code
plural_table <- matrix(
  c(18, 12,
    26,  4),
  nrow = 2,
  byrow = TRUE
)

observed_probability <- dhyper(18, m = 44, n = 16, k = 30)
fisher_result <- fisher.test(plural_table)
sample_odds_ratio <- (18 * 4) / (12 * 26)

stopifnot(abs(observed_probability - 0.01584368) < 0.00000001)
stopifnot(abs(fisher_result$p.value - 0.03909714) < 0.00000001)

c(
  observed_table_probability = observed_probability,
  two_sided_p_value = fisher_result$p.value,
  sample_odds_ratio = sample_odds_ratio
)

The two-sided p value is approximately .0391. This tail probability is larger than the probability of the observed table because it sums multiple sufficiently improbable tables.

Describing the association

The sample odds ratio for the displayed orientation is

\widehat{\operatorname{OR}} =\frac{18\times4}{12\times26} \approx.231.

The estimated odds of the productive -(e)s pattern are lower for monosyllabic lexemes than for polysyllabic lexemes. Reversing the rows or columns takes the reciprocal, so the orientation must be stated.

The odds ratio describes the sample association, while the exact test evaluates the conditional null distribution. Neither object by itself identifies why the association arises.

Stating what is conditioned on

The calculation is exact conditional on the margins: in an experiment, some margins may be fixed by design, while in a corpus sample, both margins may instead be random outcomes of the sampling process.

Fisher’s calculation conditions on the fixed margins, and that conditioning must be stated. The test can support a conditional comparison, but its reference distribution is not a complete generative account of the table.

Independence among observations also remains necessary. Sixty noun lexemes sampled from one lexical family are not equivalent to sixty independently sampled lexical types.

Association and explanation

A small p value supplies evidence against the specified independence model, but it does not establish that syllable count causes plural-pattern choice. Phonotactics, lexical frequency, etymological class, and sampling restrictions could produce the table. Distinguishing those accounts requires theoretically motivated predictors or a design that separates them.

Check your understanding

  1. Why does the upper-left count determine the other three counts when the margins are fixed?
  2. What distribution assigns probabilities to possible upper-left counts?
  3. Why is the observed-table probability not the two-sided p value?
  4. What linguistic claim remains unsupported after rejecting independence?

When a large-sample approximation is appropriate, the chi-squared test uses an approximate reference distribution rather than enumerating conditional table probabilities.