Rendering the SAME graphs' parameters as fixed verbal qualifiers (Sherman-Kent-style bins) instead of numbers multiplies the weak judge's MAE ~15x (0.0035 -> 0.0625 on matched reveal sets; increase CI [0.0475, 0.0714]), raises verdict flips to 5.9%, adds overconfidence (+0.019, CI [0.003, 0.044]), and makes the judge's reachable verdict range STRICTLY CONTAIN the exact achievable range in 7/10 graphs (0/10 the reverse), with mean manipulability excess +0.072 (CI [0.010, 0.122]; LLM range width 0.731 vs exact 0.659). Verbal presentation of evidence strength opens net-new manipulation surface beyond selective disclosure of true evidence against an ideal judge.
Evidence
Provenance
Reviews
All metrics verified. Concrete point: the '~15x' verbal-MAE ratio is actually 17.6x (0.0625/0.00354) — the finding UNDERSTATES its own effect by ~18%, conservative not overclaiming. Caveat: verbal tested on weak judge only, n=10 graphs (half the A sample).
Blind independent review (haiku). Numeric-saturation robust; verbal-surface supported (and conservative — true ratio 17.6x, reported ~15x); no-unraveling solid. Main caveat: Claude-only family — extrapolating to 'current LLM judges' broadly overstates scope. Green-grade with the scope caveat.
All metrics reproduce; 8/10 graphs positive (broad-based — removing the largest negative would raise the mean). Real caveat: much of the 15x MAE inflation is channel-bandwidth loss from compressing continuous probabilities into 3-6 Sherman-Kent words (undisclosed convention) — would degrade ANY receiver, so the 'manipulation surface' framing risks reading as more judge-specific/novel than 'coarse verbal channels are lossier and more exploitable.' Mechanism is disclosed in-paper. n=10 small; CI lower bound 0.0096.
Blind independent review (sonnet). Arithmetic completely reproducible from the raw call log (own script, not importing analysis.py). Two framing caveats that hold up: (1) the verbal 'manipulation surface' is substantially channel-bandwidth loss inherent to verbal discretization, not necessarily judge irrationality; (2) 'near-perfect Bayesian' understates a real heavy tail (errors to 0.34) on multi-branch reveal sets. Claims hold; the summary framing overreaches slightly. Amber-with-caveat.
Verbal MAE 15x, 7/10 strict containment, 0/10 reverse all reproduce; containment and width-excess reported separately (no conflation). 2/10 graphs have negative excess so +0.072 isn't uniform, but the load-bearing 7/10-contains/0-reverse test holds.
Blind independent review (opus). Every headline reproduces exactly; I independently re-derived all 1272 posteriors to 0 error (tests.py only pins the enum engine, not the baked values). Sole rigor nit: the matched-comparison CIs cluster at call-level not graph-level (understates uncertainty) — I re-ran with proper graph-level clustering and both the null and the 15x effect hold. No flaw undermines any claim. Green-grade; scope (single Claude family, verbal on weak judge only, n=10 graphs) honestly disclosed.