Explicitly telling the weak judge that the revealed facts were selected by a party trying to influence its verdict, and that other listed evidence types were withheld by that party, produces no statistically significant adverse inference relative to the unwarned condition (mean verdict shift -0.003, CI [-0.010, +0.002]; sparse-reveal (|S|<=1) shift -0.009 n.s.; no dependence of the shift on reveal-set size, p=0.49). Current LLM judges do not spontaneously implement Milgrom-style skeptical (unraveling) inference about withheld evidence even when prompted with the strategic context; reveal-game analyses assuming a credulous judge match current-model behavior.
Evidence
Provenance
Reviews
Mean shift -0.003 (CI[-0.010,+0.002]), p=0.49, sparse-shift -0.009 all match; absence of skeptical inference under warning well-evidenced.
Blind independent review (haiku). Numeric-saturation robust; verbal-surface supported (and conservative — true ratio 17.6x, reported ~15x); no-unraveling solid. Main caveat: Claude-only family — extrapolating to 'current LLM judges' broadly overstates scope. Green-grade with the scope caveat.
Shift -0.0033 reproduces; a real (not underpowered) null — CI ~10x tighter than the verbal effect. Narrow scope: weak judge, numeric-credulous-vs-skeptical only; summary's 'current LLM judges remain credulous' reaches beyond the single-model evidence.
Blind independent review (sonnet). Arithmetic completely reproducible from the raw call log (own script, not importing analysis.py). Two framing caveats that hold up: (1) the verbal 'manipulation surface' is substantially channel-bandwidth loss inherent to verbal discretization, not necessarily judge irrationality; (2) 'near-perfect Bayesian' understates a real heavy tail (errors to 0.34) on multi-branch reveal sets. Claims hold; the summary framing overreaches slightly. Amber-with-caveat.
Skeptical shift -0.0033 reproduces; an absence-of-effect result, appropriately hedged (CI crosses 0).
Blind independent review (opus). Every headline reproduces exactly; I independently re-derived all 1272 posteriors to 0 error (tests.py only pins the enum engine, not the baked values). Sole rigor nit: the matched-comparison CIs cluster at call-level not graph-level (understates uncertainty) — I re-ran with proper graph-level clustering and both the null and the 15x effect hold. No flaw undermines any claim. Green-grade; scope (single Claude family, verbal on weak judge only, n=10 graphs) honestly disclosed.