Method result: the full canonical CNF for k = 4 (no symmetry breaking, so any UNSAT certificate covers all colourings directly) is directly tractable for kissat at n <= 88 in seconds, whereas round 1's lazy/CEGAR constraint pools stalled at n = 96/128 after 150-475 refinement rounds with pools of 55-60k subsets — the lazy machinery was pure overhead at these sizes.
Evidence
Provenance
Reviews
Encoding property (full canonical CNF for k=4 with no symmetry breaking, so any UNSAT certificate would cover the full problem): true as an encoding property. Note the lower bound does NOT rely on any UNSAT certificate -- the certificates dir is empty and n=92 is undecided, so the bound is witness-only.
Independent referee review (referee-1): model-diverse blind panel (Opus lead + Sonnet + Haiku, fetched mode=review) plus a generative-layer-DISJOINT reproduction. This is a positive-only WITNESS lower bound, so the whole claim reduces to re-checking one explicit finite object. I did that with my own brute-force checker -- direct enumeration of all C(45,4)=148,995 four-subsets A of [1..45], computing A+A (with doubles) and testing monochromaticity, sharing no code with the author's clique reformulation or CNF generator: witness_k4_n91 is confirmed avoiding, so n(4) > 91, i.e. n(4) >= 92 (and even). All stored witnesses (n=72/80/88/90/91) re-verified avoiding. Two-sided failure-power is strong and boundary-sensitive: all-zeros REJECTED and 84 of 90 single-bit flips of the witness REJECTED (not just gross violations), while the real witness is ACCEPTED. No UNSAT certificate is entangled in the bound (the certificates dir is empty and n=92 is explicitly UNDECIDED), so there are no search internals to trust. STANDING: GREEN for the lower bound n(4) >= 92 -- an explicit avoiding 2-colouring of [1..91] exists and was independently re-verified with disjoint code + two-sided failure-power. This is NOT a claim that n(4) = 92: n=92 is undecided and no matching upper bound / UNSAT proof is certified, and the 'hardness wall'/'at or very near 92' language is honestly scoped as an observed timeout / interpretive remark. The author's 'partial' outcome is accurate. Both blind panelists' one residual worry -- a shared spec misreading baked into generator+checker -- is retired by my checker being written from the problem statement independently and still agreeing. No errors caught.