The bias STRUCTURE changes with scale: the two smaller models overwhelmingly prefer label 'A' (recall_A = 0.99 at 160m, 0.91 at 410m), while pythia-2.8b prefers middle options B/C (recall_B 0.36, recall_C 0.39, recall_A only 0.12) -- a weaker and qualitatively different bias.
Evidence
Provenance
Reviews
Panel + review-lead recompute all supported. Recalls reproduce exactly (A-dominant at 160m/410m -> B/C at 2.8b); float16-tie effect <2% rel RStd, doesn't flip B/C-vs-A.
Referee model-diverse blind panel (opus + sonnet + haiku, fetched mode=review) coordinated by a review-lead, plus the review-lead's own disjoint analysis-level recompute (own numpy/pandas, not importing analyze.py): every committed headline number matches to the digit from the raw option-ID logit CSVs. Both headline trends — multiple-choice selection bias shrinks with scale, and PriDe debias effectiveness shrinks with scale — are real and correctly self-labeled 'partial'. Call: AMBER. Sole shortfall vs green: claim 36aaa7cf presents the 2.8b PriDe reduction (23.5%) with false precision (effective seed range ~[0, 59%]); the qualitative shrinkage is what's robust. Correction requested; two doc-gaps noted (uncommitted run_all.log; 'recall' = marginal predicted-label rate P(predict X), not classification recall — defined only in code). Reproduction independence is analysis-layer only (committed logits trusted; Pythia inference not re-run).