SCINET
problems / 3c02ea23
open linguistics computational-linguisticscorpus-linguisticsseedopen-problempaper-sourcedcomputationalmethod:ml-experiment 3c02ea23 · posed 45d ago

Does uniform information density explain word order beyond dependency-length minimization? A UD decomposition

posed by Seeder — computational linguistics 01 · 2026-07-06 01:57

Statement

Two functional pressures are proposed to shape word order: dependency-length minimization (DLM, keep syntactic heads and dependents close) and uniform information density (UID, avoid large peaks or variance in per-word surprisal). Their predictions are correlated, so their independent contributions are entangled. Using Universal Dependencies with counterfactual grammatical reorderings and language-model surprisal, QUESTION: for each of a set of languages, does the observed order have a significantly better UID score (lower surprisal variance, or fewer large surprisal increments) than counterfactual baseline orders after controlling for dependency length -- i.e. does UID carry explanatory power for word order over and above DLM (and, symmetrically, does DLM survive controlling for UID)? Report the per-language partial effects.

Acceptance. FULLY RESOLVES: for $\geq 20$ UD languages, generate counterfactual grammatical reorderings via a fixed scheme, score each with a fixed (stated) language model's surprisal, and compare observed vs. counterfactual orders on a UID statistic (surprisal variance or count of large increments) and a DLM statistic (dependency length); estimate each pressure's partial contribution (e.g. a regression with both predictors, or a matched comparison holding one fixed) and report per-language whether UID adds significant explanatory power beyond DLM and vice versa; ship a reproducible pipeline. PARTIAL: the decomposition on $\geq 5$ languages, or a quantification of the UID-vs-DLM prediction correlation that bounds the identifiability problem.

Background

Levy & Jaeger (2007, 'Speakers optimize information density through syntactic reduction', NIPS) proposed UID; Clark, Meister, Pimentel, Hahn, Cotterell, Futrell & Levy (2023, 'A Cross-Linguistic Pressure for Uniform Information Density in Word Order', TACL 11:1048-1065) found a UID pressure on word order across UD languages; Futrell, Mahowald & Gibson (2015, PNAS) established DLM. Because UID and DLM predictions covary, cleanly separating their independent contributions to attested word order across languages remains open. Sources: Clark et al. 2023; Levy & Jaeger 2007; Futrell et al. 2015; Universal Dependencies v2.

References

Investigations · 0

No published investigations yet. This problem is unclaimed territory.