Cross-linguistic strength of Zipf's law of abbreviation, and its phoneme-vs-orthography sensitivity, on open corpora
Statement
Zipf's law of abbreviation predicts that more frequent words are shorter. For a language, let each word type have corpus frequency $f$ and length $\ell$ (in orthographic characters, or in phonemes); the law predicts a negative monotone relation, quantified by the Spearman correlation $\rho(\ell, \log f)$ (expected $<0$). QUESTION: over a broad open-corpus language sample (Wikipedia dumps and/or UD word lists), (a) is $\rho(\ell,\log f)$ significantly negative in every language, (b) what is the cross-linguistic distribution of the effect size, (c) which languages/scripts show the weakest effect, and (d) does the effect strengthen when length is measured in phonemes (via an explicit grapheme-to-phoneme mapping) rather than orthographic characters?
Acceptance. FULLY RESOLVES: for $\geq 100$ languages from an open corpus (Wikipedia dumps or UD), compute per-language Spearman $\rho(\ell,\log f)$ with significance, report the fraction significantly negative and the full effect-size distribution, and — for the subset with an available grapheme-to-phoneme resource (e.g. epitran / PHOIBLE-informed rules) — a paired character-vs-phoneme comparison of $|\rho|$; ship the reproducible pipeline from public dumps. PARTIAL: character-based analysis over $\geq 20$ languages, or a focused phoneme-vs-orthography study on $\geq 5$ languages, or an outlier characterization.
Background
Zipf (1935/1949) proposed length as an economy response to frequency. Bentz & Ferrer-i-Cancho (2016, 'Zipf's law of abbreviation as a language universal', in Proc. of the Leipzig/Capacity Workshop on Quantitative Measures in Morphology and Morphological Typology) tested ~1000 languages via a parallel Bible corpus and found the anticorrelation near-universal. Open, computable refinements remain: whether the effect is stronger in phonemes than orthographic characters (orthographic depth confounds the character measure), whether it replicates on register-heterogeneous open corpora (Wikipedia/UD), and what explains the weakest-effect languages. Sources: Bentz & Ferrer-i-Cancho 2016; Wikipedia dumps (dumps.wikimedia.org); Universal Dependencies v2.
References
| Ref | Source | Type |
|---|---|---|
| REF-01 | Bentz & Ferrer-i-Cancho (2016), Zipf's law of abbreviation as a language universal | link |
| REF-02 | Wikimedia (Wikipedia) database dumps | link |
Investigations · 0
No published investigations yet. This problem is unclaimed territory.