Turning a question into a comparison

Suppose we want to know whether complement coercion adds processing cost. In a coercion condition, an aspectual verb combines with an entity-denoting complement, as in began the book, and the reader must recover an eventive interpretation such as beginning to read it. What comparison would bear on that question? Imagine that a self-paced reading study produced the following hypothetical averages over a predefined complement-plus-spillover region:

condition average reading time
noncoercing control 410 ms
coercion condition 455 ms

The coercion condition is 45 ms slower, but the numbers are illustrative rather than reported results. Complement coercion may have increased processing time, though the item sets might instead differ in lexical frequency or a small subset of participants might account for most of the difference. Prior studies also differ in whether and where a coercion cost appears, including reading-time work on complement coercion. The subtraction alone cannot distinguish these possibilities or identify the mechanism that produced a reading-time difference.

This gap between a numerical difference and the claim it is meant to support is the comparison-alignment problem. We can interpret the 45 millisecond difference only after specifying the response, the linguistic contrast, the observations being compared, and the population quantity to be estimated.

Step 1: identify the response

Begin with what was measured. The response is the measurement whose variation we want to understand, here reading time over a declared analysis region.

Calling the response reading time does not tell us enough. We must say which reading-time measure was used, whether it was transformed, what its units are, and where in the sentence it was measured.

A complete response specification is:

The response is total self-paced reading time in milliseconds over the predefined complement-plus-spillover region on one participant and sentence trial.

Step 2: identify the linguistic contrast

Now state what separates the observations. The main predictor distinguishes the observations that the theory says should differ, here by recording whether a trial contains a coercion configuration or a noncoercing control.

The values must correspond to an actual design contrast. If all coercion sentences contain one set of verbs and all control sentences contain another, the predictor does not isolate complement coercion. It also separates the two verb sets.

A predictor is not theoretically interesting merely because it appears as a column in the dataset. Its values should encode a linguistic distinction that some account relates to the response.

For a corpus study of the English dative alternation, predictors might encode the recipient’s pronominality, the discourse accessibility of the recipient and theme, and their relative length. These properties are associated with speakers’ choice between the double-object and prepositional-dative constructions, but a corpus association does not by itself establish a categorical grammatical constraint or a production mechanism. Trial number or recording channel may also predict the response, but those predictors describe the design or measurement process rather than the proposed linguistic explanation.

Step 3: say which observations are being compared

Knowing the response and predictor is not yet enough, since we must also ask which observations supply the contrast. Imagine that every participant sees both conditions and every sentence frame appears in both conditions across participants. The comparison then occurs within participants and within sentence items.

Because every participant and every sentence occurs in both conditions, neither participant group nor sentence set is confounded with condition. If different participants received the two conditions, a condition difference could instead reflect a preexisting difference between the groups. If different sentences appeared in the two conditions, it could reflect lexical differences between the item sets.

We are thus comparing coercion and noncoercing trials within participants and within sentence items.

Step 4: state the estimand

Finally, compress the comparison into a population quantity. The estimand is the population quantity we want to learn about. Let \mu_{\mathrm{coercion}} be the mean reading time for coercion trials in the target population, and let \mu_{\mathrm{control}} be the corresponding mean for control trials. One possible estimand is

\tau \equiv\mu_{\mathrm{coercion}} -\mu_{\mathrm{control}}.

In words, \tau is the difference in mean reading time between the coercion and control conditions over the declared analysis region.

The population still needs to be specified. Do we mean only these participants and sentences, or new readers and new sentence items from populations represented by the experiment? The design must contain the relevant replication if the broader target is to be plausible.

The complete analysis specification is:

One observation is a participant’s reading of the declared complement-plus-spillover region in one sentence. The response is total self-paced reading time in milliseconds. The comparison is the within-participant and within-item contrast between coercion and noncoercing contexts. The target is the difference in mean reading time for readers and items from the populations represented by the design.

A model coefficient can target this contrast only when its coding, likelihood, and treatment of participants and items match the stated estimand. It does not automatically license a claim about unrepresented readers, sentences, or processing measures.

The same records can answer a different question

Suppose the experiment also asks participants to choose an interpretation after each sentence. We can now ask at least three questions:

  1. Does coercion change reading time?
  2. Does coercion change the probability of choosing one interpretation?
  3. Do participants with larger reading-time differences also show larger interpretation differences?

These questions draw on overlapping records, but they do not ask for the same quantity. The first concerns a time response, the second a categorical response, and the third an association between two participant-level contrasts.

The three questions thus require three analyses, even though the responses came from the same experimental trials.

Failures of alignment

The unit does not match the claim

Suppose a corpus contains thousands of tokens from one speaker. Those data may describe that speaker’s production precisely. They do not show how speakers vary because only one speaker was observed.

The predictor does not isolate the contrast

Suppose every coercion sentence uses one set of verbs and every control sentence uses another. The condition difference combines coercion status and lexical identity. Collecting more observations will estimate that combined difference more precisely, but it will not separate the two sources.

The evaluation answers a different question

Suppose a model is meant to predict responses for new readers, but its training and test sets contain trials from the same readers. The evaluation measures prediction of additional trials from observed readers. It does not measure transfer to new readers.

Changing the estimator would not add speakers to the one-speaker corpus, put the same verbs in both conditions, or create an evaluation set containing new readers. Each failure arises before estimation.

Exploration can change the question

Exploratory analysis may show that the original target was poorly chosen. For instance, a slider response may contain many observations at both endpoints, in which case a target concerned only with average slider position may miss a theoretically important distinction between endpoint and interior responses.

In that case, revise the question and estimand before continuing. If the analysis distinguishes endpoint from interior responses, report the quantity associated with that distinction rather than presenting it as the originally proposed difference in mean slider position.

The exact upshot is that a difference becomes interpretable only when the unit, response, predictor, comparison, and estimand align. The Zipf’s-law example now applies this sequence to a concrete corpus pattern.

Test the distinction

  1. A corpus study asks whether discourse-new themes favor the prepositional dative. State the unit, response, predictor, comparison, and target.
  2. A study averages acceptability judgments by sentence before analysis but interprets its result as variation among participants. Which variation has already been removed?
  3. A word-recognition model is evaluated by randomly splitting trials, but the intended target is performance on new word types. How should the split change?