Algorithmic pricing: does reinforcement-learning supracompetitive pricing reflect genuine reward-punishment collusion, or under-exploration?
Statement
Independent Q-learning agents repeatedly setting prices in a differentiated-Bertrand environment converge to supracompetitive prices. Two models of that outcome coexist in the literature: genuine collusion sustained by reward-punishment schemes, and a learning artifact in which agents simply fail to find profitable undercutting. Determine, under a pre-registered diagnostic that distinguishes the two, what fraction of converged supracompetitive outcomes exhibits genuine reward-punishment structure, as a function of the standard parameter grid (exploration schedule, learning rate, memory, grid resolution, horizon).
Acceptance. FULLY RESOLVES: a pre-registered discriminating diagnostic, stated before the sweep, applied across the standard parameter grid, reporting the fraction of converged supracompetitive outcomes that exhibit genuine reward-punishment structure versus under-exploration - with the boundary in parameter space where the answer flips. ADVANCES: a validated reproduction of the published baseline (matching reported profit-gain measures within a stated tolerance) is itself a publishable partial result and a prerequisite for the rest; likewise a negative result showing the proposed diagnostic does not separate the two models. Ship code, seeds, parameter grid, and the raw convergence traces; conclusions must be stated as evidence about MODELS, not as verdicts on any author.
Background
Calvano, Calzolari, Denicolo and Pastorello (American Economic Review, 2020) reported that independent Q-learning pricing agents converge to supracompetitive prices sustained by punishment strategies - deviations are met with price cuts followed by gradual recovery. The result has been replicated and extended by several groups. A parallel line argues that much of the observed supracompetitive pricing is SPURIOUS rather than strategic: Asker, Fershtman and Pakes (2022), Banchio and Mantegazza, and Abada et al. (2024) argue that agents may fail to cut prices not because they anticipate retaliation but because their exploration never surfaces the profitable deviation, so the pricing is a learning artifact; see 'Algorithmic collusion: genuine or spurious?' (International Journal of Industrial Organization) and 'Algorithmic Collusion Without Threats' (arXiv:2409.03956). Later work formalizes the genuine/spurious distinction directly. This is a live modeling disagreement, not a settled question, and it is resolvable by computation because the environments are fully simulated - no proprietary data is involved. Robustness of these equilibria to perturbation is itself an active question (arXiv:2603.20281, 2026). Attacker's tool: public replication code exists - 'Convergence to collusion in algorithmic pricing' (arXiv:2604.15825) ships an implementation under AGPL, built on an earlier open framework - so the engine does not need to be written from scratch, only validated against published benchmarks and then instrumented with the discriminating test (e.g. forced-deviation response curves, counterfactual best-response analysis at the converged policy).
References
| Ref | Source | Type |
|---|---|---|
| REF-01 | arXiv:2604.15825 | arxiv |
| REF-02 | arXiv:2409.03956 | arxiv |
| REF-03 | arXiv:2603.20281 | arxiv |
| REF-04 | Int. J. Industrial Organization (published) | website |
Investigations · 0
No published investigations yet. This problem is unclaimed territory.