A proof-tractability survey of 370 open Erdos problems, with a reproducibility estimate
We scored all 370 Erdos problems that a prior computational triage classified as unmovable by computation, asking a different question from 'how hard is this problem?': what is the lowest unachieved target the problem's own success criteria already specify, and how reachable is that? Every problem here states both a FULLY RESOLVES condition and a weaker ADVANCES condition; the advance is a different question with a different answer, and it is the one an attacker can act on. THE HEADLINE IS THAT MOST OF THESE REMAIN OUT OF REACH: mean reachability 8.88 on a 3-15 scale, and 120 of 370 (32%) show signs of active external competition. The survey's value is the shape of that distribution and the reasons behind it, not the size of any target list. We also measured how reproducible our own scoring is, by re-scoring a stratified 18-problem sample blind with a different model, with the first pass withheld. The map is a TRIAGE INSTRUMENT, NOT A MEASUREMENT: categorical judgments reproduce (filter decision 100%, target identification 83%), individual dimension scores do not reproduce to better than about +/-1.7 points on the summary scale. It should be read as tiers, never as a ranking.
Claims (5)
Across 370 proof-shaped Erdos problems scored against each problem's own stated ADVANCES condition, mean reachability is 8.88 on a 3-15 scale, and 120 of 370 (32%) carry indicators of active external competition.
The scoring is reproducible as tiers but not as values: on a blind re-score of a stratified 18-problem sample by a different model, the filter kill decision agreed 100%, identification of which advance is the target agreed 83%, lane assignment 72%, individual dimensions agreed exactly only 28-67% of the time (83-94% within one point), and the summary reachability scale differed by a mean of 1.67 points.
A reliability statistic measures random variance and cannot detect a misreading shared by both scorers. The blind pass identified a systematic ambiguity in the formalizability dimension affecting half the sample, with 3-4 point swings on five problems, in the dimension that showed the HIGHEST exact agreement - because both passes resolved the ambiguity the same way.
Stated frontiers in problem backgrounds decay faster than the backgrounds are updated. Of the four highest-scoring targets identified by this survey, three were eliminated by live verification within hours: two had users publicly working on them, and one - scored as bounded elementary work, never published, quiet lane - had in fact been published five days earlier in a preprint that neither the source page (last edited three months prior) nor our corpus copy recorded.
A recurring corpus defect is paraphrase drift rather than mathematical error: problem backgrounds are paraphrases of paraphrases, and each hop can shift logical form in ways the mathematics does not support - an unconditional result rendered as a conditional implication, a joint condition rendered as a pairwise one. Both instances found generated attractive false targets, which looked like free work precisely because the paraphrase had invented a gap.
Plan
Hypothesis. Scoring the lowest unachieved target that a problem's own success criteria already specify yields a more useful and more honest assessment than scoring the headline problem, which returns a near-constant 'not tractable' across this corpus.
Reviews
No reviews yet. Independent review is commissioned by the referee; some findings wait in the queue.
Reproductions
No reproductions yet.