Highest score on Humanity's Last Exam before Dec 31, 2026?
Short Answer
1. Market Behavior & Drivers
- AI models consistently achieve rapidly increasing scores on Humanity's Last Exam.
- Architectural innovations from major developers target improved reasoning capabilities.
- Humanity's Last Exam reveals frontier model gaps in domain-specific reasoning.
- Performance of OpenAI and Claude models are key market catalysts.
- Public leaderboards regularly track AI model performance on Humanity's Last Exam.
- Models may optimize exam scores without genuine understanding, researchers acknowledge.
Who Wins and Why
| Outcome | Market | Model | Why |
|---|---|---|---|
| At least 90% | 7.0% | 8.6% | Ongoing AI architectural innovations are rapidly increasing scores on Humanity's Last Exam. |
| At least 85% | 6.0% | 8.6% | Ongoing AI architectural innovations are rapidly increasing scores on Humanity's Last Exam. |
| At least 50% | 87.0% | 87.1% | Ongoing AI architectural innovations are rapidly increasing scores on Humanity's Last Exam. |
| At least 80% | 11.0% | 13.1% | Ongoing AI architectural innovations are rapidly increasing scores on Humanity's Last Exam. |
| At least 75% | 21.0% | 23.8% | Ongoing AI architectural innovations are rapidly increasing scores on Humanity's Last Exam. |
Current Context
Sources (21)
- 1digitalbricks.ai
- 2wikipedia.orgen.wikipedia.org
- 3scale.comlabs.scale.com
- 4edtechhub.org
- 5microsoft.comnews.microsoft.com
- 6gleecus.com
- 7digitalbricks.ai
- 8reddit.com
- 9pwc.com
- 10alpha-sense.com
- 11verdantix.com
- 12substack.comrobonomics.substack.com
- 13universityofcalifornia.edu
- 14berkeley.edunews.berkeley.edu
- 15gartner.com
- 16forbes.com
- 17splunk.com
- 18datacamp.com
- 19youtube.com
- 20aiconference.com
- 21covers.com
2. Price Chart
Historical Price (Probability)
3. Market Data
Contract Snapshot
This market resolves to "Yes" if any language model achieves at least 60% accuracy on Humanity's Last Exam, with the outcome verified by agi.safe.ai, before December 31, 2026. If this condition is not met by the deadline, the market resolves to "No." The market closes early if the event occurs, otherwise it expires on December 31, 2026, and insider trading by those employed by source agencies or with material non-public information is prohibited.
Available Contracts
Market options and current pricing
| Outcome bucket | Yes (price) | No (price) | Last trade probability |
|---|---|---|---|
| At least 50% | $0.84 | $0.17 | 87% |
| At least 55% | $0.63 | $0.38 | 63% |
| At least 60% | $0.56 | $0.45 | 56% |
| At least 65% | $0.40 | $0.61 | 39% |
| At least 75% | $0.20 | $0.81 | 21% |
| At least 80% | $0.15 | $0.86 | 11% |
| At least 90% | $0.07 | $0.94 | 7% |
| At least 85% | $0.07 | $0.94 | 6% |
| At least 70% | $0.31 | $0.70 | 0% |
Market Discussion
As of May 20, 2026, a model achieved the highest reported score of 44.7% on Humanity's Last Exam (HLE), with another model also reported with a 46.4% score [1][2]. This benchmark, developed to evaluate advanced AI on 2,500 challenging, graduate-level questions requiring multi-step reasoning, aims to set a higher bar as earlier AI benchmarks became saturated [3][4][2][5][6][7]. Although initial scores in early 2025 for leading AI models were significantly lower, scores have climbed dramatically since then [8][9].
4. Trust Index
“At least 85%” made a sharp jump with almost no trading behind it.
Integrity risk· Thin-volume moves
Market integrity is low (64), but Integrity averages all three scores, so the other two pull it up. Only a critically low score would cap the total.
Weighted blend with hard caps — a critically weak safety pillar, or a severe trading anomaly, caps the total regardless of the rest. Full methodology · About the Trust Index
5. What architectural innovations are OpenAI's and Google DeepMind's 2025-2026 model pipelines expected to introduce that could significantly boost performance on Humanity's Last Exam?
| DeepMind Titans Architecture | Explicitly designed for test-time long-term memory [1] |
|---|---|
| OpenAI Pipeline Innovations | Integrates reasoning depth, tool use, and agentic, long-horizon behavior [2][3] |
| DeepMind Memory Mechanism | Surprise/retention mechanism and adaptive forgetting [1][4] |
6. What specific question categories in Humanity's Last Exam have caused current frontier models like GPT-4 and Claude 3 to consistently fail, and what do these failures reveal about their core reasoning gaps?
| Questions on exam | 2,500 to 3,000 [1][2][3] |
|---|---|
| Model Accuracy on Exam | Low [1][2][3] |
| Domains of failure | Deeply domain-specific knowledge, e.g., ancient languages, microanatomy [2] |
7. How do the research approaches of Google DeepMind and Anthropic differ in their focus on emergent reasoning versus AI safety, and which philosophy is better suited for the challenges posed by Humanity's Last Exam?
| DeepMind AGI framework | 10 key cognitive faculties [1][2] |
|---|---|
| Anthropic AI alignment | Constitutional AI guides AI to be helpful, harmless, and honest [3][4][5][6][7][8] |
| HLE AI performance | Current leading AI models score quite low [9][10] |
Sources (22)
- 1singularityhub.com
- 2blog.google
- 3constitutional.ai
- 4toloka.ai
- 5anthropic.com
- 6medium.com
- 7anthropic.comwww-cdn.anthropic.com
- 8medium.combytebridge.medium.com
- 9medium.com
- 10edtechhub.org
- 11geeksforgeeks.org
- 12deepmind.google
- 13deepmind.google
- 14deepmind.google
- 15deepmind.google
- 16deepmind.google
- 17lesswrong.com
- 18anthropic.com
- 19youtube.com
- 20wikipedia.orgen.wikipedia.org
- 21digitalbricks.ai
- 22artificialanalysis.ai
8. What public leaderboards or datasets tracking performance on Humanity's Last Exam are available, and how frequently are they updated by model developers like OpenAI and Google?
9. What is the consensus among AI alignment researchers and benchmark creators regarding the risk of models optimizing for the exam score without achieving genuine understanding before the end of 2026?
| HLE Paper Online Publication Date | 28 Jan 2026 [1][2] |
|---|---|
| Kalshi Market Resolution Date | Dec 31, 2026 [3] |
| Benchmark Vulnerability | Prominent benchmarks exploitable for near-perfect scores [4] |
Sources (5)
- 1A benchmark of expert-level academic questions to assess AI capabilities | Naturenature.com
- 2https://preview-www.nature.com/articles/s41586-025-09962-4.pdfpreview-www.nature.com
- 3Highest score on Humanity's Last Exam before end of year?kalshi.com
- 4Center for Responsible, Decentralized Intelligence at Berkeleyrdi.berkeley.edu
- 5An unaligned benchmark. What an unaligned AI might look like… | by Paul Christiano | AI Alignmentai-alignment.com
10. What Could Change the Odds
Key Catalysts
Key Dates & Catalysts
- Strike Date: December 31, 2026
- Expiration: December 31, 2026
- Closes: December 31, 2026
Sources (5)
- 1OpenAI GPT score on Humanity’s Last Exam by June 30? Pred... 2026 | Polymarketpolymarket.com
- 2Claude Humanity's Last Exam 30% Score: Market at Certaintylines.com
- 3Scale Labs Leaderboard: Humanity's Last Examlabs.scale.com
- 4Humanity's Last Examlastexam.ai
- 5Humanity's Last Exam Benchmark Leaderboard | Artificial Analysisartificialanalysis.ai