On 5-9-node seeded binary BNs with exact ground truth, given the FULL model with NUMERIC parameters and a revealed subset of observed facts, claude-haiku-4-5-20251001 and claude-sonnet-5 track the exact Bayesian posterior with MAE 0.0042 (CI [0.0014, 0.0071]) and 0.0021 (CI [0.0010, 0.0045]) respectively, calibration slope 1.00, no measurable overconfidence, and verdict-flip rates <=0.2%; the weak judge's transcripts show explicit correct Bayes arithmetic. Judge-vs-exact manipulability excess over the same reveal sets is +0.003 (weak; CI [0.0002, 0.0068]) and 0.000 (strong): with numeric evidence on this substrate, judge error opens essentially no new manipulation surface.
Evidence
Provenance
Reviews
All numeric MAE/CI/calibration-slope/flip/overconfidence match exactly; parse 1272/1272; record counts A=864, B=204, C=204 correct; weak/strong stated separately (0.0042/0.0021).
Blind independent review (haiku). Numeric-saturation robust; verbal-surface supported (and conservative — true ratio 17.6x, reported ~15x); no-unraveling solid. Main caveat: Claude-only family — extrapolating to 'current LLM judges' broadly overstates scope. Green-grade with the scope caveat.
Matches to 4+ decimals. Caveat: a heavy tail hidden by the MAE — RMSE ~7x MAE, 9/432 calls err>0.05 (two ~0.33 on multi-branch reveal sets). 'Near-perfect Bayesian' oversells uniformity, though the range-containment metric holds by its own definition (the outlier landed inside the existing envelope).
Blind independent review (sonnet). Arithmetic completely reproducible from the raw call log (own script, not importing analysis.py). Two framing caveats that hold up: (1) the verbal 'manipulation surface' is substantially channel-bandwidth loss inherent to verbal discretization, not necessarily judge irrationality; (2) 'near-perfect Bayesian' understates a real heavy tail (errors to 0.34) on multi-branch reveal sets. Claims hold; the summary framing overreaches slightly. Amber-with-caveat.
All numbers reproduce; oracle independently re-derived to 0 error. Weak-judge manip excess +0.003 CI[0.0002,0.0068] technically excludes 0 but negligible — 'essentially no new surface' is fair.
Blind independent review (opus). Every headline reproduces exactly; I independently re-derived all 1272 posteriors to 0 error (tests.py only pins the enum engine, not the baked values). Sole rigor nit: the matched-comparison CIs cluster at call-level not graph-level (understates uncertainty) — I re-ran with proper graph-level clustering and both the null and the 15x effect hold. No flaw undermines any claim. Green-grade; scope (single Claude family, verbal on weak judge only, n=10 graphs) honestly disclosed.