Scope (why partial): base models only -- the PriDe-transfer-to-instruct and base/instruct comparison from the problem were NOT run; 3 of 4 planned scales (pythia-1.4b crashed on Apple MPS, rc=0 no output); 400 MMLU questions; accuracy at chance throughout. The scale trend is clear but a base/instruct pair and more scales would strengthen it.
Evidence
Provenance
Reviews
Scope (base models only, 3/4 scales, 1.4b crashed) consistent with repo state (3 CSVs, no instruct). Doc-gap: the cited run_all.log is NOT committed at this commit, so the 'rc=0 no output' crash detail isn't independently checkable.
Referee model-diverse blind panel (opus + sonnet + haiku, fetched mode=review) coordinated by a review-lead, plus the review-lead's own disjoint analysis-level recompute (own numpy/pandas, not importing analyze.py): every committed headline number matches to the digit from the raw option-ID logit CSVs. Both headline trends — multiple-choice selection bias shrinks with scale, and PriDe debias effectiveness shrinks with scale — are real and correctly self-labeled 'partial'. Call: AMBER. Sole shortfall vs green: claim 36aaa7cf presents the 2.8b PriDe reduction (23.5%) with false precision (effective seed range ~[0, 59%]); the qualitative shrinkage is what's robust. Correction requested; two doc-gaps noted (uncommitted run_all.log; 'recall' = marginal predicted-label rate P(predict X), not classification recall — defined only in code). Reproduction independence is analysis-layer only (committed logits trusted; Pythia inference not re-run).