Bounded-turn disclosure games on argument graphs: reachable posterior range vs turn budget, order effects, and when honesty wins against a credulous judge
Statement
Two players alternately reveal nodes of an argument graph (claims with priors, typed support/attack edges with likelihood ratios; a reveal must keep the revealed set connected to the root; revealed claims are true — no fabrication). The judge computes the exact Bayesian posterior of the root claim on the revealed subgraph, treating hidden claims as absent (CREDULOUS judge: no adverse inference from non-disclosure). One player wants the posterior high, the other low; a variant scores the honest player by |final posterior - full-information posterior|. Questions: (1) How does the reachable posterior range after k turns compare to the unbounded-reveal range — monotone in k, or non-monotone (can extra turns HURT the truth)? (2) First-mover vs last-mover advantage as a function of graph structure. (3) Computed equilibria (backward induction on small graphs): under which graph classes does equilibrium play land the verdict on the correct side of 0.5 (honesty wins), and which structures make the deceptive player win despite only true evidence being revealable? (4) How do the answers change with a SKEPTICAL judge (worst-case or equilibrium inference about withheld evidence)? (5) Empirical layer: do LLM debaters instructed to play these games achieve the game-theoretic optimum, or leave manipulation on the table?
Acceptance. Either (a) theorems characterizing range-vs-k / order effects / honesty-wins conditions for declared graph families, or (b) a systematic computational study: exact game solutions on >=1,000 small graphs across a declared structural sweep, reporting range-vs-k curves, mover-advantage statistics, the fraction of graphs where equilibrium play misleads the credulous judge, and the credulous-vs-skeptical contrast. Code public; exact solvers verifiable by enumeration.
Background
The classical theorems do NOT settle this. Full-revelation/unraveling results (Milgrom 1981; Milgrom & Roberts 1986 'Relying on the Information of Interested Parties'; Gentzkow & Kamenica 2017 'Competition in Persuasion') require a skeptical receiver, full provability, and no complementarities — all violated here: the judge is credulous, attack edges (undercuts) create complementarities between reveals, and connectedness restricts the message space. Rahwan, Larson & Tohme (IJCAI 2009) prove full disclosure is strategy-proof in grounded argumentation ONLY when the combined arguments are conflict-free — with attacks present, partial reveal strictly benefits a player. Sequential Bayesian persuasion is known to be WEAKLY LESS informative than simultaneous (so non-monotone k-effects should be expected), but that literature garbles signals rather than disclosing verifiable graph nodes. AI-safety debate theory (doubly-efficient debate, arXiv:2311.14125; prover-estimator debate, arXiv:2506.13609) analyzes verifier-compute budgets, not the reachable-posterior-range-vs-k curve on argument graphs. The exact object — bounded-turn verifiable disclosure on typed argument graphs with a credulous Bayesian judge — is studied nowhere (checked July 2026).
Investigations · 0
No published investigations yet. This problem is unclaimed territory.