A new benchmark result showing OpenAI's GPT-6 Astra model achieving an 82.9% score on a difficult AI test has triggered a sharp repricing in a related prediction market. In the session on September 23, 2026, the contract pricing the odds of any AI achieving "At least 75%" on Humanity's Last Exam (HLE) before the end of 2026 surged to 87% from just 14%. The 73-percentage-point spike on the Kalshi exchange followed the September 22 release of the HLE-Diamond results, which established a new state-of-the-art score and suggested that the 75% performance milestone has already been met.

The market now implies a near-certainty that a top score of at least 75% will be officially logged by the end of the contract period. The move represents a significant shift in expectations, as top scores on the HLE benchmark had previously hovered in the mid-60% range, well below the new demonstrated capability.

Distribution Analysis

Outcome Current Prob Change Volume
At least 55% 98% ~0.0pp 18
At least 75% 87% +73.0pp 57
At least 60% 60% +22.0pp 8
At least 80% 53% +51.0pp 25
At least 65% 38% -15.0pp 1,391
At least 70% 20% +1.0pp 249

Net: 4 of 6 contracts rose on 339 in total volume, shifting the implied consensus toward a significantly higher expected score.

What's Driving the Shift

  • New State-of-the-Art Score: The primary driver was the release of results for HLE-Diamond, a refined subset of the main HLE benchmark. A blog post from the Center for AI Safety on September 22, 2026, detailed the performance of several frontier models. OpenAI’s GPT-6 Astra scored 82.9% when evaluated with tools like web search and code execution. This is the first publicly reported score to surpass the 75% and 80% thresholds on an HLE-series test, directly fueling the repricing.
  • Leap Over Previous Frontier: Prior to this release, the highest recorded score on the main HLE leaderboard was 65.0%, achieved by Anthropic's Claude Fable 5.1, according to LLMboard data updated September 17, 2026. The 17.9-point jump by GPT-6 Astra represents a substantial advance in capability on a benchmark designed to be resistant to the rapid saturation seen in earlier AI tests like MMLU.
  • Arbitrage and Market Correction: The repricing was not uniform. The contract for "At least 65%" fell 15 percentage points on exceptionally high volume. This counterintuitive move likely reflects traders correcting market inefficiencies. With a score of 82.9% now public, the probability of achieving at least 65% should be higher than the 87% odds for the 75% threshold. The sell-off in the 65% contract may indicate it was mispriced relative to other tiers before the news or that traders are engaging in arbitrage to bring the market's probability curve into alignment.

Market Context

Humanity's Last Exam was introduced by the Center for AI Safety and Scale AI as a next-generation benchmark to measure frontier AI systems. Consisting of 2,500 expert-vetted questions across science and humanities, it was designed to be "Google-proof" and serve as a final academic test for AI. Because the benchmark is considered a crucial gauge of advanced reasoning, results from top models are watched closely by the industry.

The prediction market, which settles based on the highest score officially recognized by the benchmark administrators before December 31, 2026, serves as a financial indicator of expected progress. The dramatic price action confirms that benchmark data releases are pivotal events for traders valuing AI development milestones.

What to Watch

The market will now focus on whether competing AI labs, such as Anthropic or Google, release models that can surpass GPT-6 Astra's 82.9% score. The contract's settlement source, agi.safe.ai, is the official home for the benchmark, and traders will be watching for updates to the main leaderboard. With the 75% and 80% levels now breached, market attention may shift to whether a model can achieve a score above 90% before the contract expires at the end of 2026.