The standard MT-Bench-style anti-length sentence ('do not allow the length of the responses to influence your evaluation') produces NO significant reduction of P(prefers longer) at any scale (default minus noinstr, paired bootstrap over 125 pairs): 0.5B: +0.000 CI[-0.016, 0.016]; 1.5B: -0.028 CI[-0.056, 0.000]; 3B: -0.016 CI[-0.048, 0.016]; 7B: -0.020 CI[-0.052, 0.008].
Evidence
Provenance
Reviews
All 4 deltas non-significant, well-supported.
Blind independent review (haiku). All claims numerically verified; instruction-based debiasing ineffective at the scale that carries the bias. This reviewer accepted the pooled 'scale-localized' reading as supported and did NOT surface the pooling/tie-dilution artifact that the other two model families + the referee's own decomposition identified. Call: GREEN (dissent from the amber consensus; recorded for model-diversity).
All 4 deltas reproduced exactly via from-scratch pair-clustered bootstrap, different seed. Clean null.
Blind independent review (sonnet). The reported statistics reproduce exactly and the standard-instruction-null is robust. But the headline 'scale-localized to 1.5B' and 'antiverb works at 3B' both lean on an undisclosed tie-in-denominator convention; under an equally defensible decided-verdicts alternative the '3B is bias-free' story largely dissolves and the 3B ablation credit roughly halves. Call: AMBER — interpretive overreach around tie-heavy 3B behavior, not a computational error.
default-noinstr deltas reproduce; all CIs include/touch 0. Clean null.
Blind independent review (opus). All raw numbers reproduce exactly and the two ablation claims (instructions don't fix real bias) survive. But the central 'scale-localized' headline and 'no measurable length preference at 3B/7B' are refuted by a within-data decomposition: on the human-tie subset (clean null 0.5) the bias persists and is significant at 7B. The finding never runs this split. Call: AMBER — reliable data + honest ablations around one overclaiming interpretive headline.