Turning a question into a comparison

Suppose we want to know whether an aspectual mismatch makes a sentence harder to process. We design two conditions and observe the following average reading times at a critical verb:

condition average reading time
aspectually matched 410 ms
aspectually mismatched 455 ms

The mismatch condition is 45 ms slower, but we do not yet know what produced the difference. Aspectual mismatch might have increased processing time. The mismatch items might instead contain lower-frequency verbs, or a small subset of participants might account for most of the difference. The subtraction alone cannot distinguish these possibilities.

We can interpret the 45 millisecond difference only after specifying the response, the linguistic contrast, the observations being compared, and the population quantity to be estimated.

Step 1: identify the response

The response is the measurement whose variation we want to understand. In this experiment, it might be reading time at the critical verb.

Calling the response reading time does not tell us enough. We must say which reading-time measure was used, whether it was transformed, what its units are, and where in the sentence it was measured.

A complete response specification is:

The response is the total reading time in milliseconds at the critical verb on one participant and sentence trial.

Step 2: identify the linguistic contrast

The main predictor distinguishes the observations that the theory says should differ. Here it records whether a trial contains an aspectually matched or mismatched context.

The values must correspond to an actual design contrast. If all matched sentences contain one set of verbs and all mismatched sentences contain another, the predictor does not isolate aspectual mismatch. It also separates the two verb sets.

A predictor is not theoretically interesting merely because it appears as a column in the dataset. Its values should encode a linguistic distinction that some account relates to the response.

For a dative-choice study, predictors might encode whether the recipient is pronominal, whether the theme is discourse-new, and how the lengths of the two arguments compare. Pronominality, discourse status, and relative argument length test proposals about accessibility, information structure, and production difficulty. Trial number or recording channel may also predict the response, but those predictors describe the design or measurement process rather than the proposed linguistic explanation.

Step 3: say which observations are being compared

Imagine that every participant sees both conditions and every sentence appears in both conditions across participants. The comparison then occurs within participants and within sentence items.

Because every participant and every sentence occurs in both conditions, neither participant group nor sentence set is confounded with condition. If different participants received the two conditions, a condition difference could instead reflect a preexisting difference between the groups. If different sentences appeared in the two conditions, it could reflect lexical differences between the item sets.

We are thus comparing aspectually matched and mismatched trials within participants and within sentence items.

Step 4: state the estimand

The estimand is the population quantity we want to learn about. Let \mu_{\mathrm{mismatch}} be the mean reading time for mismatch trials in the target population, and let \mu_{\mathrm{match}} be the corresponding mean for match trials. One possible estimand is

\tau \equiv\mu_{\mathrm{mismatch}} -\mu_{\mathrm{match}}.

In words, \tau is the difference in mean reading time between the mismatch and match conditions.

The population still needs to be specified. Do we mean only these participants and sentences, or new readers and new sentence items from populations represented by the experiment? The design must contain the relevant replication if the broader target is to be plausible.

The complete analysis specification is:

One observation is a participant’s reading of the critical verb in one sentence. The response is total reading time in milliseconds. The comparison is the within-participant and within-item contrast between aspectually matched and mismatched contexts. The target is the difference in mean reading time for readers and items from the populations represented by the design.

A model coefficient can target this contrast only when its coding, likelihood, and treatment of participants and items match the stated estimand. It does not automatically license a claim about unrepresented readers, sentences, or processing measures.

The same records can answer a different question

Suppose the experiment also asks participants to choose an interpretation after each sentence. We can now ask at least three questions:

  1. Does mismatch change reading time?
  2. Does mismatch change the probability of choosing one interpretation?
  3. Do participants with larger reading-time differences also show larger interpretation differences?

These questions draw on overlapping records, but they do not ask for the same quantity. The first concerns a time response, the second a categorical response, and the third an association between two participant-level contrasts.

The three questions thus require three analyses, even though the responses came from the same experimental trials.

Failures of alignment

The unit does not match the claim

Suppose a corpus contains thousands of tokens from one speaker. Those data may describe that speaker’s production precisely. They do not show how speakers vary because only one speaker was observed.

The predictor does not isolate the contrast

Suppose every mismatch sentence uses one set of verbs and every match sentence uses another. The condition difference combines mismatch and lexical identity. Collecting more observations will estimate that combined difference more precisely, but it will not separate the two sources.

The evaluation answers a different question

Suppose a model is meant to predict responses for new readers, but its training and test sets contain trials from the same readers. The evaluation measures prediction of additional trials from observed readers. It does not measure transfer to new readers.

Changing the estimator would not add speakers to the one-speaker corpus, put the same verbs in both conditions, or create an evaluation set containing new readers. Each failure arises before estimation.

Exploration can change the question

Exploratory analysis may show that the original target was poorly chosen. For instance, a slider response may contain many observations at both endpoints, in which case a target concerned only with average slider position may miss a theoretically important distinction between endpoint and interior responses.

In that case, revise the question and estimand before continuing. If the analysis distinguishes endpoint from interior responses, report the quantity associated with that distinction rather than presenting it as the originally proposed difference in mean slider position.

Test the distinction

  1. A corpus study asks whether discourse-new themes favor the prepositional dative. State the unit, response, predictor, comparison, and target.
  2. A study averages acceptability judgments by sentence before analysis but interprets its result as variation among participants. Which variation has already been removed?
  3. A word-recognition model is evaluated by randomly splitting trials, but the intended target is performance on new word types. How should the split change?