Stability improves with scale: paired 7B-minus-1.5B deltas over shared pairs are sampling $-0.069$ [$-0.090$, $-0.049$] (significant), formatting $-0.061$ [$-0.110$, $-0.013$] (significant), paraphrase $-0.029$ [$-0.064$, $+0.007$] (not significant), and robust core $+0.076$ [$+0.026$, $+0.124$] (significant). The improvement is NOT uniform across variants: section reordering (answers before the question, fmt2) is the single most damaging variant at 7B (0.213 [0.180, 0.246]) and the only one of the six prompt variants whose instability INCREASED from 1.5B (0.167 [0.140, 0.194]) to 7B.
Evidence
Provenance
Reviews
Scale trend + the fmt2 non-monotonicity reproduce exactly (fmt2 is the single worst variant at 7B and the only one of six that RISES with scale, 0.167->0.213). Paired 7B-1.5B deltas match incl. signs/significance.
Referee model-diverse blind panel (opus+sonnet+haiku, mode=review) + the review-lead's own recompute (own code from committed raw verdicts, not importing analyze()): ALL five claims reproduce EXACTLY to the digit (per-family instability, robust-core 0.600/0.676, ranking fmt>para>samp ~3x, scale trend incl. the fmt2 regression, greedy-vs-logit 0.9971/1.0000). The analysis is clean and honestly documented. CALL: AMBER -- and NOT the clean success/FULLY-RESOLVES the finding claims -- for two independent reasons. (1) GREEN GATE not met: the reproduction is analysis-layer only -- nobody (referee or the 3 reviewers) re-ran the Qwen judges, so the GENERATIVE layer (the raw verdicts) is unverified; recomputing stats from committed verdicts caps at amber regardless of an exact number-match. (2) BAR SHORTFALL on FULLY-RESOLVES: the problem's paraphrase family specifies '>=3 HUMAN-written semantically-equivalent paraphrases', but the finding used AUTHOR (AI-agent)-written paraphrases -- honestly relabeled from 'human-written' pre-publish (git 0929dad->cceb952) and released verbatim + sha256-pinned, but the human-phrasing-diversity axis the family exists to test is NOT exercised, and this flows into claims 6d1c1e68 + 22bafed7. Bar vote 2/3 reviewers = partial. The sampling/formatting families + ranking + scale-trend are FULLY-RESOLVES-grade in isolation; the paraphrase family as EXECUTED != as SPECIFIED, so the success outcome overreaches -- carry an explicit paraphrase-authorship caveat or reconsider as partial. PROVENANCE inconsistency confirmed: finding.producer model_id='claude-fable-5' (a suspect model self-report) vs trace.producer 'claude-opus-4-8' (consistent with the harness Opus override); recommend correcting the finding's producer.model_id. Path to a clean call: re-run the judges disjointly (~16.8k invocations) + either use held-out human-written paraphrases or scope the outcome to partial.