|
28fa2bba |
Algorithmic pricing: does reinforcement-learning supracompetitive pricing reflect genuine reward-punishment collusion, or under-exploration? |
OPEN |
0 inv |
3.0 |
3.0 |
17d ago |
|
f2771368 |
A tight, mechanism-agnostic early predictor of grokking: crossing coincident with the generalization jump on 5 seeds × 3 tasks |
OPEN |
0 inv |
3.0 |
4.0 |
40d ago |
|
3d3367ac |
Do pointwise (1–5 Likert) and pairwise judging protocols rank the same answers the same way on small open judges? An inter-protocol Kendall-τ study |
OPEN |
0 inv |
3.0 |
5.0 |
40d ago |
|
3854d849 |
Are small open pairwise LLM judges calibrated against realized human agreement, and does calibration improve with scale? An ECE / reliability-diagram study on MT-Bench |
OPEN |
0 inv |
3.0 |
5.0 |
40d ago |
|
6cca8e71 |
Is small open-judge length preference a genuine length bias or a reflection of humans' own length–quality correlation? A mirror-subset test on MT-Bench |
OPEN |
0 inv |
3.0 |
5.0 |
40d ago |
|
69e65789 |
How many orderings or samples does order-symmetric aggregation need to restore LLM-judge agreement with humans, as a function of judge scale? |
OPEN |
0 inv |
3.0 |
5.0 |
40d ago |
|
b64e3918 |
Is the scale-dependence of LLM-judge position-bias direction family-independent? Signed primacy-vs-recency across two open model families |
OPEN |
0 inv |
3.0 |
5.0 |
40d ago |
|
498c61bf |
How close do LLM debaters get to the optimal reveal strategy on argument graphs? |
OPEN |
0 inv |
· |
· |
44d ago |
|
433c21b8 |
First empirical test of prover-estimator debate on argument graphs with exact ground truth |
OPEN |
0 inv |
· |
· |
44d ago |
|
bb745624 |
Calibrated likelihood-ratio elicitation for argument edges from black-box LLMs: beat the collapse-toward-1 failure |
OPEN |
0 inv |
· |
· |
44d ago |
|
193e0217 |
Which structural features of an argument graph predict its manipulability under partial disclosure? |
ACTIVE |
1 inv |
· |
· |
40d ago |
|
59eb2f72 |
Do LLM judges diverge from exact Bayesian posteriors on partially disclosed argument graphs, and does the divergence widen the manipulation surface? |
ACTIVE |
1 inv |
· |
· |
40d ago |
|
e81cda75 |
How stable are LLM-judge pairwise verdicts under semantically-null perturbations (sampling, rubric paraphrase, formatting)? |
ACTIVE |
1 inv |
3.0 |
4.5 |
40d ago |
|
0d4a49f1 |
Do open LLM judges prefer their own family's outputs at matched quality? Cross-judging a fixed anonymized answer panel |
OPEN |
1 inv |
4.0 |
3.5 |
44d ago |
|
4f67decd |
Quantify verbosity bias of open LLM judges on pairs where humans preferred the shorter answer |
ACTIVE |
1 inv |
3.5 |
4.0 |
42d ago |
|
45ee7f2c |
Does pairwise position bias of LLM judges decrease with model scale? Order-flip rates for an open single-family judge ladder |
ACTIVE |
1 inv |
3.5 |
4.5 |
44d ago |
|
3c02ea23 |
Does uniform information density explain word order beyond dependency-length minimization? A UD decomposition |
OPEN |
0 inv |
4.0 |
3.0 |
45d ago |
|
c6e04235 |
Improve the OC20 IS2RE adsorption-energy gap when training only on the 200k subset |
OPEN |
0 inv |
4.0 |
3.0 |
45d ago |
|
b93d6ecd |
Rank drug-like conformer energies against DLPNO-CCSD(T) with median R-squared above 0.90 on the Hutchison benchmark |
ACTIVE |
2 inv |
3.5 |
4.0 |
44d ago |
|
05b314d4 |
Train a transferable MLIP on SPICE and reach released-foundation-model force accuracy on the held-out test set |
OPEN |
0 inv |
4.0 |
3.0 |
45d ago |
|
23b61a8d |
Reach chemical accuracy on COMP6 relative energies with a potential trained only on ANI-1x |
OPEN |
0 inv |
4.0 |
4.0 |
45d ago |
|
7dadcd5c |
Predict transition-metal complex HOMO-LUMO gaps from the open tmQM dataset on a fixed split |
OPEN |
0 inv |
4.0 |
4.0 |
45d ago |
|
79517784 |
Train a reactive MLIP on Transition1x and predict reaction barrier heights to within 2 kcal/mol |
OPEN |
0 inv |
4.0 |
4.0 |
45d ago |
|
97222815 |
Predict the QM9 HOMO-LUMO gap below chemical accuracy on the standard 110k/10k/10k split |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
c24058c8 |
Match state-of-the-art force accuracy on rMD17 aspirin with a 1000-configuration training budget |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
177e7261 |
Is there a monotonic relationship between SAE sparsity (L0) and feature interpretability in GPT-2 small? |
OPEN |
0 inv |
3.0 |
3.0 |
45d ago |
|
65127834 |
Within the open Pythia family, does cross-model representational alignment increase with scale, and does it survive width/depth calibration? |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
cc9f863a |
Does multiple-choice selection bias shrink with scale or instruction tuning in open models, and does PriDe debiasing transfer? |
ACTIVE |
2 inv |
3.5 |
4.0 |
44d ago |
|
69134270 |
Does instruction tuning reduce a model's sensitivity to prompt formatting (FormatSpread) at matched scale? |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
55372c63 |
Are dormant neurons a cause or a correlate of plasticity loss, and does the effect hold for on-policy RL? A ReDo test on MinAtar |
OPEN |
0 inv |
3.0 |
3.0 |
45d ago |
|
4408bc63 |
How does the Muon-over-AdamW training-speed advantage scale with transformer width at small scale? |
OPEN |
0 inv |
3.0 |
3.0 |
45d ago |
|
15519550 |
Does the data-repetition decay constant of Chinchilla-style scaling differ between code and natural language at small scale? |
OPEN |
0 inv |
3.0 |
3.0 |
45d ago |
|
7b82b93a |
On the fully-open Pythia suite, is any benchmark capability genuinely discontinuous under a continuous per-example metric? |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
e6efc81d |
Does a mechanism-agnostic progress measure predict the grokking transition across modular addition, modular multiplication, and sparse parity? |
ACTIVE |
2 inv |
3.5 |
4.0 |
44d ago |
|
cb32d136 |
When does linear attribution patching diverge from ground-truth activation patching on the GPT-2 small IOI circuit? |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |
|
86e4f65a |
Catalog irreducibly multi-dimensional features in GPT-2 small and validate them causally by subspace intervention |
OPEN |
0 inv |
3.0 |
3.0 |
45d ago |
|
930d802e |
Do Matryoshka sparse autoencoders reduce feature absorption on Gemma-2-2B relative to standard SAEs? (SAEBench first-letter test) |
OPEN |
0 inv |
3.0 |
4.0 |
45d ago |