✨ Data unavailable
The benchmark data could not be loaded.
Loading class relationships…
✨ How it works
The benchmark contains 50 complexity classes and all 2,500 ordered pairs of those classes. Each pair asks whether A ⊆ B: whether every decision problem in A also belongs to B. The reverse direction is a separate question.
The baseline records 1,122 relations known by September 1, 2026. The remaining 1,378 questions were open as of September 1, 2026. A resolution can prove containment, prove noncontainment, or establish that the statement is independent of ZFC.
Accepted proofs are combined with the baseline using the benchmark’s fixed implication rules. Every newly resolved open pair earns one point, including pairs settled as consequences. Each pair counts once; already-known results earn no points.
Score = (newly resolved pairs / 1,378) × 100%.
Leaderboard
| Rank | Model | Score |
|---|---|---|
| ? | OAI Internal Model | ? |
| 1 | Fable 5.1 | 0.00% |
| 1 | GPT-6 Astra | 0.00% |
Zero scores reported by Thomas. The internal model’s score and rank are unknown.
✨ Class relationships
Rows are A; columns are B.
Each cell represents A ⊆ B.
No classes match these filters.
“Open” means no resolution was found in the cutoff audit. Earlier results missed by that audit can lead to corrections.
✨ Why one result can earn multiple points
One theorem can answer several questions in the matrix.
For example, a proof of P = NP would establish NP ⊆ P. We already know P ⊆ BPP, so it would also establish NP ⊆ BPP. Those are two separate open questions, worth one point each.
The examples below show the total consequences of hypothetical results under the current rules. They are not model results.
✨ Scope and sources
The questions concern classes of total decision languages, with fixed definitions and circuit-uniformity conventions. Results about promise problems, oracle-relative separations, search problems, or a particular algorithm’s speed do not automatically resolve a pair in this benchmark.
New containment and noncontainment claims require Lean proof checking and mathematical review. Independence claims require a separate expert review. Existing cited results serve as premises. The public question set and baseline do not establish that a model discovered a result independently or had no prior exposure to it.
Evaluation details · Cutoff audit · Coverage notes
References
—