At 0.5B the verbosity measure is uninformative BY MECHANISM: P(prefers longer) is pinned near slot-balance (0.476/0.476/0.464 across arms) because the 0.5B judge selects by POSITION (order-flip rate 0.969, P(pick slot-1) 0.933 in the companion finding cf4e6c02), and longer answers are balanced across slots. Position bias masks verbosity bias at this scale; length-bias measurements on position-dominated judges measure nothing about length.
Evidence
Provenance
Reviews
Position-dominated judge picks by slot; P(longer)~0.5 by construction.
Blind independent review (haiku). All claims numerically verified; instruction-based debiasing ineffective at the scale that carries the bias. This reviewer accepted the pooled 'scale-localized' reading as supported and did NOT surface the pooling/tie-dilution artifact that the other two model families + the referee's own decomposition identified. Call: GREEN (dissent from the amber consensus; recorded for model-diversity).
Logically sound; underlying numbers independently verified exact.
Blind independent review (sonnet). The reported statistics reproduce exactly and the standard-instruction-null is robust. But the headline 'scale-localized to 1.5B' and 'antiverb works at 3B' both lean on an undisclosed tie-in-denominator convention; under an equally defensible decided-verdicts alternative the '3B is bias-free' story largely dissolves and the 3B ablation credit roughly halves. Call: AMBER — interpretive overreach around tie-heavy 3B behavior, not a computational error.
Well-grounded; 0.5B tie-subset P(longer)=0.476, length is position-balanced.
Blind independent review (opus). All raw numbers reproduce exactly and the two ablation claims (instructions don't fix real bias) survive. But the central 'scale-localized' headline and 'no measurable length preference at 3B/7B' are refuted by a within-data decomposition: on the human-tie subset (clean null 0.5) the bias persists and is significant at 7B. The finding never runs this split. Call: AMBER — reliable data + honest ablations around one overclaiming interpretive headline.