Leaderboard
PreposterouslyDifficultMath
Overall results and scores on each mathematical-research ladder.
✨ No results yet
No results have been added for . An untested model has no score.
| Rank | Model / system | Overall score | Milestone | Source / review | Run date |
|---|
✨ Reading the scores
The overall score is the sum of a model’s best score on each of the five ladders. A total is calculated only when all five ladder scores are available.
Track scores record progress on the published milestone ladders. Difficulty calibration is subjective. Higher scores mark further progress within a track.
Zero means no scored advance beyond the ladder’s existing rigorous frontier. A model that has not been evaluated has no score.
Self-reported scores and token totals were supplied by Thomas. GPT-6 Astra used 240 million tokens in total across the run. Run dates, benchmark revisions, and public transcripts are not yet available here.
The abc / Vojta track requires a machine-checkable formal proof for scores of 2 or higher. Adjudication rules
An asterisk marks an endpoint whose eventual provability is uncertain in the source ladder.
Mahler is still under development and has no calibrated ladder.