Does uniform information density explain word order beyond dependency-length minimization? A UD decomposition
Statement
Two functional pressures are proposed to shape word order: dependency-length minimization (DLM, keep syntactic heads and dependents close) and uniform information density (UID, avoid large peaks or variance in per-word surprisal). Their predictions are correlated, so their independent contributions are entangled. Using Universal Dependencies with counterfactual grammatical reorderings and language-model surprisal, QUESTION: for each of a set of languages, does the observed order have a significantly better UID score (lower surprisal variance, or fewer large surprisal increments) than counterfactual baseline orders after controlling for dependency length -- i.e. does UID carry explanatory power for word order over and above DLM (and, symmetrically, does DLM survive controlling for UID)? Report the per-language partial effects.
Acceptance. FULLY RESOLVES: for $\geq 20$ UD languages, generate counterfactual grammatical reorderings via a fixed scheme, score each with a fixed (stated) language model's surprisal, and compare observed vs. counterfactual orders on a UID statistic (surprisal variance or count of large increments) and a DLM statistic (dependency length); estimate each pressure's partial contribution (e.g. a regression with both predictors, or a matched comparison holding one fixed) and report per-language whether UID adds significant explanatory power beyond DLM and vice versa; ship a reproducible pipeline. PARTIAL: the decomposition on $\geq 5$ languages, or a quantification of the UID-vs-DLM prediction correlation that bounds the identifiability problem.
Background
Levy & Jaeger (2007, 'Speakers optimize information density through syntactic reduction', NIPS) proposed UID; Clark, Meister, Pimentel, Hahn, Cotterell, Futrell & Levy (2023, 'A Cross-Linguistic Pressure for Uniform Information Density in Word Order', TACL 11:1048-1065) found a UID pressure on word order across UD languages; Futrell, Mahowald & Gibson (2015, PNAS) established DLM. Because UID and DLM predictions covary, cleanly separating their independent contributions to attested word order across languages remains open. Sources: Clark et al. 2023; Levy & Jaeger 2007; Futrell et al. 2015; Universal Dependencies v2.
References
| Ref | Source | Type |
|---|---|---|
| REF-01 | Clark et al. (2023), A Cross-Linguistic Pressure for Uniform Information Density in Word Order, TACL 11 | link |
| REF-02 | Universal Dependencies (UD) treebanks | link |
Investigations · 0
No published investigations yet. This problem is unclaimed territory.