Short Answer

The market considers ChatGPT likely to be the top coding AI this month (56.0%), diverging from the model's 40.8% probability.

1. Market Behavior & Drivers

The 11.0 percentage point increase on September 2 coincides with an Anthropic announcement for two new models, Claude Fable 5.1 and Mythos 5.1. This release appears to have lifted trader expectations for Claude's performance on the DeepSWE benchmark.
The larger 19.0 percentage point spike on September 8 lacks a clear catalyst in the provided materials. The supplied context for that date discusses a significant price drop in a related market, which is inconsistent with the positive price action observed here. Overall market sentiment is informed by benchmark data from early September showing Claude Opus 5 at a 73.6% pass rate, narrowly trailing the leader, Gemini 3.8 Flash, which scored between 73.8% and 74%.
  • ChatGPT leads; GPT-6 Astra scored 74.1% on DeepSWE in early September.
  • Claude and Gemini remain strong from prior DeepSWE v1.1 leaderboard performance.
  • New 'Agentic Search' tool may influence final DeepSWE scores for leading models.

Who Wins and Why

Outcome Market Model Why
Gemini 29.0% 29.2% Gemini shows strong potential, integrating deeply with Google's ecosystem and tools.
ChatGPT 56.0% 40.8% ChatGPT leads in coding capabilities, benefiting from frequent updates and continuous user feedback.
Claude 35.0% 23.0% Claude is rapidly improving with its long context window, enhancing suitability for complex coding tasks.
Grok 1.0% 1.0% Grok offers real-time information access, but its coding prowess remains unproven against leaders.
GLM 2.0% 2.0% GLM models are robust for specific programming languages, yet lack broad adoption.

Current Context

Current benchmarks show close competition; Anthropic is favored in predictions. As of early September 2026, the DeepSWE coding benchmark, which comprises 113 original, long-horizon software engineering tasks designed to evaluate coding agents [^], is led by Gemini 3.8 Flash at a 73.8%74% pass rate [^][^]. Claude Opus 5 follows closely at 73.6%, and GPT-5.6 Sol at 72.7% [^]. Meta's Muse Spark 1.3 has reportedly achieved 75.4% in non-official evaluations [^]. OpenAI introduced GPT-6 Astra in September 2026, a frontier model optimized for autonomous agentic computer use, which achieved a 72.6% success rate on the OSWorld 2.0 benchmark. Early expert feedback notes strong practical utility for engineering, though large projects still require substantial human coordination [^][^][^]. Despite the current DeepSWE leaders, prediction markets as of early September 2026 heavily favor Anthropic (79–93% implied probability) to hold the top-ranked AI model by month's end, reflecting consensus on their recent Claude series performance [^][^][^].
Agentic coding methodologies advance rapidly, alongside industry consolidation. Recent developments in agentic coding include 'Harness-of-Harness' for multi-day autonomous development [^], 'ACToR' for adaptive retrieval [^], 'Speculative Uncertainty' for failure detection [^], and 'KVMem' for managing long-context workspaces [^]. Entire launched "Agentic Search" in September 2026, a tool designed to enhance coding agents by enabling semantic context retrieval (commits, history, design decisions) across repositories, which reduces token usage and agent steps in complex tasks [^]. The DeepSWE Verifier, a trained system used to reward and evaluate coding agents during reinforcement learning, has shown improved performance with alternative reward methods such as 'Dockerless' as of September 2026 [^]. Key industry developments include Nvidia’s $12.9 billion agreement to acquire Hugging Face, signaling consolidation within the open-source AI ecosystem [^]. Further, major model updates from Anthropic and other developers are anticipated later in September [^][^].

2. Price Chart

Historical Price (Probability)

Outcome probability
Date

3. Significant Price Movements

Notable price changes detected in the chart, along with research into what caused each movement.

Outcome: Claude

📉 September 08, 2026: 59.0pp drop

Price decreased from 94.0% to 35.0%

What happened: The 59.0 percentage point drop in "Claude's" market price on September 8, 2026, appears to be a delayed market re-evaluation of previously reported significant performance regressions for Claude Code. Earlier in 2026 (February–April), traditional news and analyses widely reported a ~67% drop in Claude Code's thinking depth, associated cost spikes, and overall quality degradation [^]. This re-pricing was likely exacerbated by rival coding AIs, such as Meta's Muse Spark 1.3 and Gemini, demonstrating strong performance on the DeepSWE benchmark in early September 2026, highlighting Claude's relative decline [^]. Social media was not a primary driver of this specific movement; rather, the underlying performance issues were extensively covered by traditional news and analysis months prior.

📈 September 07, 2026: 66.0pp spike

Price increased from 28.0% to 94.0%

What happened: The primary driver of the 66.0 percentage point price spike for "Claude" in the "Top Coding AI this month (DeepSWE)" market was Anthropic's announcement on September 4, 2026, that Claude completed the first machine-verified proof of Fermat's Last Theorem using 13 million lines of Lean code [^]. This unprecedented achievement in formal mathematics and logic, requiring sophisticated code generation and verification, strongly signaled Claude's advanced capabilities as a coding AI, leading the market to re-evaluate its potential [^]. This news preceded the market movement, appearing to lead the price spike. No significant social media activity related to Claude's DeepSWE performance was reported around this date, nor was there any specific 66.0pp performance spike on DeepSWE for Claude on September 7, 2026 [^].

📈 September 03, 2026: 12.0pp spike

Price increased from 14.0% to 26.0%

What happened: The primary driver for the price spike was social media activity announcing a significant coding performance gain. Surge AI announced a 12.0 percentage point improvement on Terminal-Bench 2.0 after training a model on their coding dataset [^]. This specific 12.0pp performance gain is noted in AI discourse concerning coding agent benchmarks observed in early September 2026, coinciding with the market movement [^][^]. While Claude experienced elevated error rates across multiple models on September 3, 2026 [^][^][^][^][^], this substantial benchmark improvement for a coding AI likely bolstered market confidence in Claude's overall competitive position, leading to the spike. Social media was a primary driver.

📉 September 02, 2026: 36.0pp drop

Price decreased from 50.0% to 14.0%

What happened: The provided research does not corroborate a 36.0 percentage point drop for "Claude" related to the "DeepSWE" benchmark on September 2, 2026 [^][^][^]. While Anthropic announced the release of Claude Fable 5.1 and Mythos 5.1 on that date [^][^][^], there is no information linking this launch to a specific performance decline on DeepSWE or the reported price movement. The "36.0pp" figure appears in unrelated AI contexts, not as a catalyst for the September 2026 Claude release [^][^][^][^]. No social media activity contributing to this predicted price movement was found in the provided information.

Outcome: ChatGPT

📈 September 01, 2026: 51.0pp spike

Price increased from 19.0% to 70.0%

What happened: The requested 51.0 percentage point spike for "ChatGPT" on the DeepSWE benchmark appears to stem from a misunderstanding of reported data. Research indicates that the "51.0 pp" figure is cited in literature as a performance degradation gap for VLA models under paraphrased instructions, rather than an AI coding benchmark spike for DeepSWE [^]. While the DeepSWE leaderboard experienced rapid shifts in early September 2026, with Meta's Muse Spark 1.3 achieving a 20.4 percentage point improvement, these events were not linked to ChatGPT or a 51.0 pp increase [^]. Therefore, no primary driver can be identified for the described price movement as the premise itself is unsupported by the provided information.

4. Market Data

Contract Snapshot

This prediction market asks which AI model will be deemed the 'Top Coding AI this month (DeepSWE)'. The exact criteria for determining the 'Top Coding AI', which would trigger a YES resolution for that AI and a NO resolution for others, are not specified in the provided content. Trading begins on September 30, 12:00 AM EDT, and the market has a maximum payout date of September 30, 2026; no special settlement conditions are detailed.

Available Contracts

Market options and current pricing

Outcome bucket Yes (price) No (price) Last trade probability
ChatGPT $0.75 $0.74 56%
Claude $0.80 $0.95 35%
Gemini $0.58 $0.89 29%
GLM $0.82 $1.00 2%
Kimi $0.95 $1.00 2%
DeepSeek $0.94 $1.00 1%
Grok $0.97 $1.00 1%
MiMo $0.82 $1.00 1%

Market Discussion

As of early September 2026, Meta's Muse Spark 1.3 is reported to lead the DeepSWE benchmark at 75.4% pass@1, with Gemini 3.8 Flash (73.8%) and GPT-6 Astra (74.1%) also featuring prominently in recent snapshots [^][^][^]. However, prediction markets strongly favor Anthropic's Claude models to secure the #1 position on other major coding benchmarks like LiveBench and Code Arena by the end of September 2026 [^][^][^][^]. Developer discussions in September 2026 indicate a shift from simple code generation metrics to focus on the reliability, verification, and cost economics of autonomous coding agents, with Meta's Muse Spark 1.3 specifically noted for its efficiency improvements [^][^][^][^][^][^][^].

5. Trader Dashboard

A deterministic, per-market integrity scorecard computed from order-book and price data. Higher is better for Trader Trust, Liquidity, Move Quality and Resolution; higher means more risk for Quote Risk and Avoid Risk.

GeminiPrimaryTrader TrustLiquidityMove Quality77ResolutionQuote RiskAvoid Risk
Move Quality77Confirmedhigh confidence
  • Factor
  • Factor
flow_agreement
0.85
move_retained_pct
100
ClaudeTrader TrustLiquidityMove Quality73ResolutionQuote RiskAvoid Risk
Move Quality73Confirmedhigh confidence
  • Factor
  • Factor
flow_agreement
0.85
move_retained_pct
100
ChatGPTTrader TrustLiquidityMove Quality60ResolutionQuote RiskAvoid Risk
Move Quality60Mostly confirmedhigh confidence
  • Factor
  • Factor
flow_agreement
0.85
move_retained_pct
83
GrokTrader TrustLiquidityMove Quality79ResolutionQuote RiskAvoid Risk
Move Quality79Confirmedmedium confidence
  • Factor
  • Factor
flow_agreement
0.85
move_retained_pct
100
DeepSeekTrader TrustLiquidityMove QualityResolutionQuote RiskAvoid Risk
Move QualityInsufficient Datainsufficient confidence
  • Factor
Metric
KimiTrader TrustLiquidityMove QualityResolutionQuote RiskAvoid Risk
Move QualityInsufficient Datainsufficient confidence
  • Factor
Metric
MiMoTrader TrustLiquidityMove QualityResolutionQuote RiskAvoid Risk
Move QualityInsufficient Datainsufficient confidence
  • Factor
Metric
GLMTrader TrustLiquidityMove QualityResolutionQuote RiskAvoid Risk
Move QualityInsufficient Datainsufficient confidence
  • Factor
Metric

trader_dashboard_lean_v1.14 · computed Sep 8, 2026

6. What upcoming model releases from Anthropic, Google, or OpenAI could alter the DeepSWE leaderboard before September 30, 2026?

GPT-6 Astra DeepSWE Score74.1% [^]
DeepSWE Leaderboard UpdateSeptember 3, 2026 [^]
DeepSWE Prediction Market Target90% Pass@1 [^][^]
Recent flagship models have established current DeepSWE leaderboard standings. Major model releases occurred in early September 2026, including Gemini 3.8 Flash on September 2, Claude Fable 5.1/Mythos 5.1, and GPT-6 Astra on September 3 [^][^][^]. These models have already been incorporated into the DeepSWE leaderboard, which was updated on September 3, 2026 [^][^][^]. Currently, GPT-6 Astra leads the benchmark with a score of 74.1%, closely followed by Gemini 3.8 Flash and Claude Opus 5 [^][^][^].
DeepSWE targets a significantly higher benchmark than current model capabilities. Prediction markets for the DeepSWE benchmark are focused on models achieving a 90% Pass@1 score, with the benchmark set to resolve on September 30, 2026 [^][^]. This 90% target is considerably higher than the approximately 74% performance exhibited by the current frontier models [^][^].
No new model releases are indicated before the DeepSWE resolution. The provided information does not indicate any upcoming model releases from Anthropic, Google, or OpenAI that could alter the DeepSWE leaderboard between the September 3, 2026 update and the September 30, 2026 resolution date.

7. What technical advantages justify prediction markets pricing Anthropic's Claude as the overwhelming favorite to top the DeepSWE benchmark in September 2026?

Claude Opus 5 DeepSWE v1.1 Score68.8% [^][^]
Sep 2026 DeepSWE Prediction - Gemini 3.8 Flash73.8% [^]
Sep 2026 DeepSWE Prediction - Claude Opus 573.6% [^]
DeepSWE is a benchmark designed to evaluate frontier coding agents on diverse, long-horizon software engineering tasks. This benchmark assesses 113 original, long-horizon software engineering tasks that are contamination-free, diverse, and verified through hand-written functional tests [^][^][^]. An update to DeepSWE v1.1 improved reliability by using a clean, isolated environment for execution and scoring [^]. Anthropic's own documentation reports that Claude Opus 5 achieved a score of 68.8% on v1.1 of this benchmark [^][^].
Prediction markets show a highly competitive DeepSWE landscape for 2026, rather than any single model being an overwhelming favorite. Forecasts for September 2026 indicate a tightly competitive top of the DeepSWE leaderboard [^]. These forecasts project Gemini 3.8 Flash achieving 73.8%, Claude Opus 5 at 73.6%, and GPT-5.6 Sol at 72.7% [^]. This competitive scenario suggests that any perceived favoritism for Claude would likely reflect market sentiment regarding expected future updates or unreleased configurations, rather than static leaderboard dominance [^].
No specific technical advantages justify Claude as an overwhelming favorite in these markets. The available research does not disclose specific technical advantages that would justify prediction markets pricing Anthropic's Claude as an overwhelming favorite to top the DeepSWE benchmark in September 2026 [^][^][^]. The bound facts do not reveal any technical advantages of Anthropic's Claude that would explain such a favored position in the market forecasts [^][^][^].

8. How do Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 compare on the specific long-horizon software engineering tasks within the DeepSWE benchmark?

Top DeepSWE v1.1 Score74.0% +- 4% (Google Gemini 3.8 Flash, Anthropic Claude Opus 5) [^][^][^]
DeepSWE v1.1 Benchmark Gap9.5% (meaningful gap for 113-task corpus) [^][^][^]
Claude Opus 5 Internal DeepSWE v1.1 Score68.8% [^][^][^]
Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 lead the DeepSWE benchmark, currently tied at the top of the DeepSWE v1.1 leaderboard. Both models achieved a score of 74.0% +- 4%, positioning them as the leading models for long-horizon software engineering tasks within this benchmark [^][^][^]. However, the DeepSWE benchmark has a documented meaningful gap of 9.5% across its 113-task corpus. This substantial gap indicates that minor score variations between models frequently fall within the benchmark's inherent noise floor [^][^][^].
Vendor-reported and independent benchmark results show discrepancies, particularly for Anthropic's Claude Opus 5. Anthropic's internal testing reported a lower score of 68.8% on DeepSWE v1.1 for Claude Opus 5, which contrasts with the 74.0% score observed on the official leaderboard [^][^][^]. This difference highlights the potential variations between performance figures released by vendors and those obtained from independent community runs on the leaderboard, which often utilize standardized harness settings such as mini-swe-agent [^][^][^].

9. What is the official update schedule and data submission methodology for the DeepSWE leaderboard for the September 2026 resolution period?

DeepSWE Update SchedulePeriodically, when submitted by contributors [^][^]
DeepSWE Submission MethodContact Datacurve team (serena@datacurve.ai) [^][^]
DeepSWE Evaluation Method'mini-swe-agent' harness on Modal platform [^]
The DeepSWE leaderboard updates periodically, requiring direct submission to the Datacurve team. There is no formal, automated public update schedule; instead, new model results are added as they are submitted by contributors [^][^]. To submit data, individuals must contact the Datacurve team directly via serena@datacurve.ai, ensuring their results adhere to the established standard methodology for formatting [^][^].
DeepSWE submissions require evaluation using a specific harness on the Modal platform. The official methodology mandates that all models submitted to the DeepSWE leaderboard must be evaluated using the 'mini-swe-agent' harness [^]. This evaluation must be performed on the Modal platform to maintain consistency and comparability across all submissions [^].
DeepSWE is not the primary reference for the September 2026 coding AI market. The "Top Coding AI this month" prediction market, which resolves on September 30, 2026, primarily references the LM Code Arena Leaderboard [^]. While this particular market does not use DeepSWE, a previous market in July 2026 did utilize DeepSWE for its resolution [^].

10. How might new agentic tools like Entire's 'Agentic Search' influence the final September 2026 DeepSWE scores for leading models like Gemini and GPT-6?

Agentic Search performance improvement70/90 to 81/90 correct answers on engineering-history questions [^]
GPT-6 Astra DeepSWE pass rate74.1% (as of September 8, 2026) [^]
GPT-5.6 Sol DeepSWE pass rate72.7% [^]
Entire's new "Agentic Search" tool enhances AI coding agent capabilities. Launched in early September 2026, this specialized tool provides coding agents with both code search functionalities and crucial semantic context, including commit histories, agent session transcripts, and reasoning [^]. Entire asserts that this tool significantly boosts agent performance on engineering-history questions, citing an increase in correct answers from 70 out of 90 to 81 out of 90 [^].
DeepSWE benchmarks prevent external tools, thus limiting direct score influence. The DeepSWE v1.1 benchmark environments are designed as isolated, sandboxed systems with strictly controlled network access, preventing evaluating agents from utilizing external, dynamic tools or APIs such as Entire's "Agentic Search" [^][^]. Consequently, existing research does not specify how "Agentic Search" might directly affect the final DeepSWE scores of leading models like GPT-6 Astra [^][^].
GPT-6 Astra currently leads the DeepSWE benchmark in early September. As of September 8, 2026, GPT-6 Astra holds the top position on the DeepSWE benchmark with a 74.1% pass rate [^][^]. Other high-performing models include GPT-5.6 Sol, which achieves a 72.7% pass rate, and Claude Fable 5 [^][^].

11. What Could Change the Odds

Key Catalysts

The DeepSWE coding benchmark is a current focus, with Meta's Muse Spark 1.3 leading at 75.4% and OpenAI's GPT-6 Astra at 74.1%, as of early September 2026 [^] [^] [^] . Gemini 3.8 Flash closely follows, ranking in the 73.8% to 74% range [^]. Prediction markets are currently focused on the September 30, 2026, resolution to determine the top-ranked model [^]. Anthropic leads market sentiment, with its Claude model showing 78.8% in some markets, though OpenAI's September 3 launch of GPT-6 Astra represents a significant market catalyst [^][^][^][^]. Early September has seen a surge of major model releases, with Anthropic's mid-September developer event also anticipated [^][^].
Bullish market sentiment stems from sustained hyperscaler investment and the high density of frontier model launches in early September [^] [^] [^] . Conversely, bearish concerns include long-term sustainability, potential infrastructure bubbles, and impending end-of-year enterprise budget cycles [^][^][^]. Ongoing EU AI Act compliance deadlines also remain a factor [^][^]. Beyond static benchmarks, coding agent performance is increasingly evaluated using multi-turn user-driven benchmarks like SWE-Interact and realistic compositional evaluations such as RealSWE, with models like Qwen3-32B showing leading results through test-time scaling [^]. Additionally, the September 2026 'Hakken' system, which predicts scientific discoveries by fusing temporal knowledge graphs with LLM semantic knowledge, could act as a catalyst for future science-tech prediction markets [^].

Key Dates & Catalysts

  • Expiration: September 30, 2026
  • Closes: September 30, 2026

12. Decision-Flipping Events

  • Trigger: The DeepSWE coding benchmark is a current focus, with Meta's Muse Spark 1.3 leading at 75.4% and OpenAI's GPT-6 Astra at 74.1%, as of early September 2026 [^] [^] [^] .
  • Trigger: Gemini 3.8 Flash closely follows, ranking in the 73.8% to 74% range [^] .
  • Trigger: Prediction markets are currently focused on the September 30, 2026, resolution to determine the top-ranked model [^] .
  • Trigger: Anthropic leads market sentiment, with its Claude model showing 78.8% in some markets, though OpenAI's September 3 launch of GPT-6 Astra represents a significant market catalyst [^] [^] [^] [^] .

14. Historical Resolutions

Historical Resolutions: 16 markets in this series

Outcomes: 2 resolved YES, 14 resolved NO

Recent resolutions:

  • KXCODEAI-26AUG31-MIMO: NO (Aug 31, 2026)
  • KXCODEAI-26AUG31-KIMI: NO (Aug 31, 2026)
  • KXCODEAI-26AUG31-GROK: NO (Aug 31, 2026)
  • KXCODEAI-26AUG31-GLM: NO (Aug 31, 2026)
  • KXCODEAI-26AUG31-GEMI: NO (Aug 31, 2026)