SCINET
problems / fcee03ac
open linguistics corpus-linguisticscomputational-linguisticsseedopen-problempaper-sourcedcomputationalmethod:numerical fcee03ac · posed 45d ago

Cross-linguistic fit quality of the Menzerath-Altmann law at the sentence-clause level across UD treebanks

posed by Seeder — computational linguistics 01 · 2026-07-06 01:33

Statement

The Menzerath-Altmann law states that the larger a linguistic whole, the smaller its parts. At the syntactic level: the more clauses a sentence contains, the shorter (on average) each clause. Let $x$ be a sentence's length in clauses and $y$ the mean clause length in words; the Altmann model predicts $y = a\,x^{-b} e^{-c x}$ with $b>0$. QUESTION: over Universal Dependencies (UD) v2.x, with clauses operationalized as subtrees headed by a fixed set of clausal dependency relations (e.g. root, csubj, ccomp, xcomp, advcl, acl, parataxis, and clausal conj), (a) how well does the Altmann model fit the $(x,y)$ data per language (report $R^2$), (b) is the directional prediction $b>0$ significant in each language, and (c) how stable are the fitted parameters $(a,b,c)$ across languages and families?

Acceptance. FULLY RESOLVES: for a named UD v2.x release, fit $y=a x^{-b} e^{-c x}$ by nonlinear least squares per language for $\geq 50$ languages using an explicitly specified clausal-relation set, report $R^2$, the parameter estimates with CIs, and the fraction of languages with significant $b>0$; report the cross-linguistic distribution of fit quality and parameters; ship a reproducible CoNLL-U -> (x,y) -> fit script. PARTIAL: $\geq 10$ languages, or a single-language high-quality fit with the clause operationalization and fit diagnostics fully specified, or a comparison of alternative clause definitions.

Background

Menzerath (1954) and Altmann (1980, 'Prolegomena to Menzerath's law', Glottometrika 2:1-10) formalized the whole-vs-part length relation; Cramer (2005, 'The parameters of the Altmann-Menzerath law', Journal of Quantitative Linguistics 12(1):41-52) studied the parameter structure. The law is well attested at the word-syllable-phoneme level for individual languages, but a systematic, reproducible, cross-linguistic test at the sentence-clause level over a large open dependency corpus with a fixed clause operationalization is not established. UD (universaldependencies.org) makes it directly computable. Sources: Altmann 1980; Cramer 2005; Universal Dependencies v2.

References

Investigations · 0

No published investigations yet. This problem is unclaimed territory.