# Top Coding AI this month (DeepSWE)

On Sep 30, 2026

Updated: September 8, 2026

Category: Science and Technology

Tags: AI

HTML: /markets/science-and-technology/ai/top-coding-ai-this-month-deepswe/

## Short Answer

**Key takeaway.** The **market** considers ChatGPT likely to be the top coding AI this month (**56.0%**), diverging from the **model**'s **40.8%** **probability**.

## Key Claims (September 2026)

**- - ChatGPT leads; GPT-6 Astra scored 74.1% on DeepSWE in early September.** - Claude and Gemini remain strong from prior DeepSWE v1.1 leaderboard performance.
- New 'Agentic Search' tool may influence final DeepSWE scores for leading models.

## Market Behavior & Drivers

Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 leading the DeepSWE benchmark drives **market** probabilities near **29%**.

The 11.0 percentage point increase on September 2 coincides with an Anthropic announcement for two new models, Claude Fable 5.1 and Mythos 5.1. This release appears to have lifted trader expectations for Claude's performance on the DeepSWE benchmark.

The larger 19.0 percentage point spike on September 8 lacks a clear catalyst in the provided materials. The supplied context for that date discusses a significant price drop in a related market, which is inconsistent with the positive price action observed here. Overall market sentiment is informed by benchmark data from early September showing Claude Opus 5 at a 73.6% pass rate, narrowly trailing the leader, Gemini 3.8 Flash, which scored between 73.8% and 74%.

### Who Wins and Why

| Outcome | Market | Model | Why |
| --- | --- | --- | --- |
| Gemini | 29.0% | 29.2% | Gemini shows strong potential, integrating deeply with Google's ecosystem and tools. |
| ChatGPT | 56.0% | 40.8% | ChatGPT leads in coding capabilities, benefiting from frequent updates and continuous user feedback. |
| Claude | 35.0% | 23.0% | Claude is rapidly improving with its long context window, enhancing suitability for complex coding tasks. |

## Model vs Market

| Outcome | Market Probability | Octagon Model Probability |
| --- | --- | --- |
| Gemini | 29.0% | 29.2% |
| ChatGPT | 56.0% | 40.8% |
| Claude | 35.0% | 23.0% |
| Grok | 1.0% | 1.0% |
| Kimi | 2.0% | 2.0% |
| MiMo | 1.0% | 1.0% |
| GLM | 2.0% | 2.0% |
| DeepSeek | 1.0% | 1.0% |

- Expiration: September 30, 2026

## Significant Price Movements

### Outcome: Claude

#### 📉 September 08, 2026: 59.0pp drop

Price decreased from 94.0% to 35.0%

**What happened:** The 59.0 percentage point drop in "Claude's" market price on September 8, 2026, appears to be a delayed market re-evaluation of previously reported significant performance regressions for Claude Code. Earlier in 2026 (February–April), traditional news and analyses widely reported a ~67% drop in Claude Code's thinking depth, associated cost spikes, and overall quality degradation [[^]](https://devtoolpicks.hashnode.dev/anthropic-explains-the-claude-code-quality-drop-here-is-what-actually-happened.md). This re-pricing was likely exacerbated by rival coding AIs, such as Meta's Muse Spark 1.3 and Gemini, demonstrating strong performance on the DeepSWE benchmark in early September 2026, highlighting Claude's relative decline [[^]](https://deepswe.datacurve.ai/). Social media was not a primary driver of this specific movement; rather, the underlying performance issues were extensively covered by traditional news and analysis months prior.

#### 📈 September 07, 2026: 66.0pp spike

Price increased from 28.0% to 94.0%

**What happened:** The primary driver of the 66.0 percentage point price spike for "Claude" in the "Top Coding AI this month (DeepSWE)" market was Anthropic's announcement on September 4, 2026, that Claude completed the first machine-verified proof of Fermat's Last Theorem using 13 million lines of Lean code [[^]](https://gigazine.net/gsc_news/en/20260907-claude-fermat-last-theorem-formalizing/). This unprecedented achievement in formal mathematics and logic, requiring sophisticated code generation and verification, strongly signaled Claude's advanced capabilities as a coding AI, leading the market to re-evaluate its potential [[^]](https://gigazine.net/gsc_news/en/20260907-claude-fermat-last-theorem-formalizing/). This news preceded the market movement, appearing to lead the price spike. No significant social media activity related to Claude's DeepSWE performance was reported around this date, nor was there any specific 66.0pp performance spike on DeepSWE for Claude on September 7, 2026 [[^]](https://claude-news.today/en/briefings/briefing-2026-09-07/).

#### 📈 September 03, 2026: 12.0pp spike

Price increased from 14.0% to 26.0%

**What happened:** The primary driver for the price spike was social media activity announcing a significant coding performance gain. Surge AI announced a 12.0 percentage point improvement on Terminal-Bench 2.0 after training a model on their coding dataset [[^]](https://www.linkedin.com/posts/surge-ai_activity-7489008334247976960-LYIx). This specific 12.0pp performance gain is noted in AI discourse concerning coding agent benchmarks observed in early September 2026, coinciding with the market movement [[^]](https://www.linkedin.com/posts/surge-ai_activity-7489008334247976960-LYIx)[[^]](https://arxiv.org/abs/2609.05274). While Claude experienced elevated error rates across multiple models on September 3, 2026 [[^]](https://pulsetic.com/status/claude/incidents/6939/)[[^]](https://pulsetic.com/status/claude/incidents/6938/)[[^]](https://www.thenews.com.pk/latest/1414756-claude-ai-faces-partial-outage-what-users-need-to-know)[[^]](https://sundayguardianlive.com/trending/claude-outage-today-is-claude-ai-down-today-thousands-users-report-claude-spike-across-multiple-models-app-errors-login-issues-claude-downdetector-status-276398/)[[^]](https://tech.sportskeeda.com/laptops/news-is-claude-right-now-september-3-2026-outage-status-explored), this substantial benchmark improvement for a coding AI likely bolstered market confidence in Claude's overall competitive position, leading to the spike. Social media was a primary driver.

#### 📉 September 02, 2026: 36.0pp drop

Price decreased from 50.0% to 14.0%

**What happened:** The provided research does not corroborate a 36.0 percentage point drop for "Claude" related to the "DeepSWE" benchmark on September 2, 2026 [[^]](https://arxiv.org/pdf/2604.08571)[[^]](https://arxiv.org/html/2603.14761)[[^]](https://arxiv.org/html/2606.05228). While Anthropic announced the release of Claude Fable 5.1 and Mythos 5.1 on that date [[^]](https://sdtimes.com/claude-fable-5-1/61089/)[[^]](https://www.the1news.com/article/claude-fable-51-and-mythos-51)[[^]](https://truescho.com/en/blog/claude-fable-5-1-release-2026), there is no information linking this launch to a specific performance decline on DeepSWE or the reported price movement. The "36.0pp" figure appears in unrelated AI contexts, not as a catalyst for the September 2026 Claude release [[^]](https://orrery.me/markets/will-the-next-claude-opus-debut-at-a-score-of-at-least-1490-by-december-31-2026-20260715004553750)[[^]](https://orrery.me/markets/will-the-next-claude-opus-model-be-released-by-july-24-2026-231-543)[[^]](https://assignee.net/audit)[[^]](https://arxiv.org/html/2606.10546). No social media activity contributing to this predicted price movement was found in the provided information.

### Outcome: ChatGPT

#### 📈 September 01, 2026: 51.0pp spike

Price increased from 19.0% to 70.0%

**What happened:** The requested 51.0 percentage point spike for "ChatGPT" on the DeepSWE benchmark appears to stem from a misunderstanding of reported data. Research indicates that the "51.0 pp" figure is cited in literature as a performance degradation gap for VLA models under paraphrased instructions, rather than an AI coding benchmark spike for DeepSWE [[^]](https://arxiv.org/html/2607.27146v1). While the DeepSWE leaderboard experienced rapid shifts in early September 2026, with Meta's Muse Spark 1.3 achieving a 20.4 percentage point improvement, these events were not linked to ChatGPT or a 51.0 pp increase [[^]](https://www.lookonchain.com/feeds/71283). Therefore, no primary driver can be identified for the described price movement as the premise itself is unsupported by the provided information.

## Contract Snapshot

This prediction market asks which AI model will be deemed the 'Top Coding AI this month (DeepSWE)'. The exact criteria for determining the 'Top Coding AI', which would trigger a YES resolution for that AI and a NO resolution for others, are not specified in the provided content. Trading begins on September 30, 12:00 AM EDT, and the market has a maximum payout date of September 30, 2026; no special settlement conditions are detailed.

## Market Discussion

As of early September 2026, Meta's Muse Spark 1.3 is reported to lead the DeepSWE benchmark at 75.4% pass@1, with Gemini 3.8 Flash (73.8%) and GPT-6 Astra (74.1%) also featuring prominently in recent snapshots [[^]](https://deepswe.datacurve.ai/)[[^]](https://benchlm.ai/benchmarks/deepswe)[[^]](https://www.kucoin.com/news/flash/deepswe-rankings-shift-gemini-rises-meta-s-muse-spark-1-3-hits-75-4). However, prediction markets strongly favor Anthropic's Claude models to secure the #1 position on other major coding benchmarks like LiveBench and Code Arena by the end of September 2026 [[^]](https://www.coinrithm.com/en/prediction-markets/kalshi/kxtechranklistaicode-26sep30)[[^]](https://www.lines.com/prediction-markets/tech/which-company-has-the-best-code-arena-webdev-ai-model-end-of-september-20260717140116512)[[^]](https://predictparity.com/markets/p/will-anthropic-have-the-best-ai-model-on-livebench-coding-at-the-end-of-september-2026-20260728165808282)[[^]](https://predictparity.com/markets/p/will-openai-have-the-best-code-arena-webdev-ai-at-the-end-of-september-2026-20260717140116515). Developer discussions in September 2026 indicate a shift from simple code generation metrics to focus on the reliability, verification, and cost economics of autonomous coding agents, with Meta's Muse Spark 1.3 specifically noted for its efficiency improvements [[^]](https://discuss.ai.google.dev/t/gemini-3-5-3-8-flash-review-on-real-ml-and-ds-jobs/180901/3)[[^]](https://dev.to/monuminu/the-agentic-coding-era-is-here-how-autonomous-ai-coding-agents-are-rewriting-the-sdlc-5dpa)[[^]](https://community.openai.com/t/gpt-5-6-sol-high-reasoning-appears-broken-instant-replies-no-deep-thinking-file-workflows-failing/1394605)[[^]](https://tech-insider.org/what-muse-spark-1-3-actually-changes/)[[^]](https://arxiv.org/abs/2609.04681)[[^]](https://arxiv.org/html/2608.13730v1)[[^]](https://arxiv.org/abs/2608.30701).

## Market Data

| Contract | Yes Bid | Yes Ask | Last Price | Volume | Open Interest |
| --- | --- | --- | --- | --- | --- |
| ChatGPT | 26% | 75% | 56% | $8,074.4 | $3,587.97 |
| Claude | 5% | 80% | 35% | $4,098.69 | $2,119.68 |
| DeepSeek | 0% | 94% | 1% | $1,136.68 | $1,136.68 |
| Gemini | 11% | 58% | 29% | $8,536.14 | $3,645.88 |
| GLM | 0% | 82% | 2% | $1,138.68 | $1,138.68 |
| Grok | 0% | 97% | 1% | $2,138.35 | $1,354.09 |
| Kimi | 0% | 95% | 2% | $1,138.68 | $1,138.68 |
| MiMo | 0% | 82% | 1% | $1,138.68 | $1,138.68 |

## What upcoming model releases from Anthropic, Google, or OpenAI could alter the DeepSWE leaderboard before September 30, 2026?

GPT-6 Astra DeepSWE Score | 74.1% [[^]](https://deepswe.datacurve.ai/) |
DeepSWE Leaderboard Update | September 3, 2026 [[^]](https://deepswe.datacurve.ai/) |
DeepSWE Prediction Market Target | 90% Pass@1 [[^]](https://manifold.markets/adssx/will-deepswe-be-90-solved-at-5task)[[^]](https://manifold.markets/adssx/will-deepswe-be-saturated-before-20) |

**Recent flagship models have established current DeepSWE leaderboard standings**

Recent flagship models have established current DeepSWE leaderboard standings. Major **model** releases occurred in early September 2026, including Gemini 3.8 Flash on September 2, Claude Fable 5.1/Mythos 5.1, and GPT-6 Astra on September 3 [[^]](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/)[[^]](https://www.anthropic.com/claude-fable-and-mythos-5-1)[[^]](https://openai.com/index/safety-overview-gpt-6-astra/). These models have already been incorporated into the DeepSWE leaderboard, which was updated on September 3, 2026 [[^]](https://deepswe.datacurve.ai/)[[^]](https://benchlm.ai/benchmarks/deepswe)[[^]](https://www.datalearner.com/en/benchmarks/deepswe-v1-1-openai-2026-09). Currently, GPT-6 Astra leads the benchmark with a score of **74.1%**, closely followed by Gemini 3.8 Flash and Claude Opus 5 [[^]](https://deepswe.datacurve.ai/)[[^]](https://benchlm.ai/benchmarks/deepswe)[[^]](https://www.datalearner.com/en/benchmarks/deepswe-v1-1-openai-2026-09).

DeepSWE targets a significantly higher benchmark than current **model** capabilities. Prediction markets for the DeepSWE benchmark are focused on models achieving a **90%** Pass@1 score, with the benchmark set to resolve on September 30, 2026 [[^]](https://manifold.markets/adssx/will-deepswe-be-90-solved-at-5task)[[^]](https://manifold.markets/adssx/will-deepswe-be-saturated-before-20). This **90%** target is considerably higher than the approximately **74%** performance exhibited by the current frontier models [[^]](https://manifold.markets/adssx/will-deepswe-be-90-solved-at-5task)[[^]](https://manifold.markets/adssx/will-deepswe-be-saturated-before-20).

No new **model** releases are indicated before the DeepSWE resolution. The provided information does not indicate any upcoming **model** releases from Anthropic, Google, or OpenAI that could alter the DeepSWE leaderboard between the September 3, 2026 update and the September 30, 2026 resolution date.

## What technical advantages justify prediction markets pricing Anthropic's Claude as the overwhelming favorite to top the DeepSWE benchmark in September 2026?

Claude Opus 5 DeepSWE v1.1 Score | 68.8% [[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1)[[^]](https://github.com/datacurve-ai/deep-swe/issues/65) |
Sep 2026 DeepSWE Prediction - Gemini 3.8 Flash | 73.8% [[^]](https://benchlm.ai/benchmarks/deepswe) |
Sep 2026 DeepSWE Prediction - Claude Opus 5 | 73.6% [[^]](https://benchlm.ai/benchmarks/deepswe) |

**DeepSWE is a benchmark designed to evaluate frontier coding agents on diverse, long-horizon software engineering tasks**

DeepSWE is a benchmark designed to evaluate frontier coding agents on diverse, long-horizon software engineering tasks. This benchmark assesses 113 original, long-horizon software engineering tasks that are contamination-free, diverse, and verified through hand-written functional tests [[^]](https://deepswe.datacurve.ai/)[[^]](https://arxiv.org/html/2607.07946)[[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1). An update to DeepSWE v1.1 improved reliability by using a clean, isolated environment for execution and scoring [[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1). Anthropic's own documentation reports that Claude Opus 5 achieved a score of **68.8%** on v1.1 of this benchmark [[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1)[[^]](https://github.com/datacurve-ai/deep-swe/issues/65).

Prediction markets show a highly competitive DeepSWE landscape for 2026, rather than any single **model** being an overwhelming favorite. Forecasts for September 2026 indicate a tightly competitive top of the DeepSWE leaderboard [[^]](https://benchlm.ai/benchmarks/deepswe). These forecasts project Gemini 3.8 Flash achieving **73.8%**, Claude Opus 5 at **73.6%**, and GPT-5.6 Sol at **72.7%** [[^]](https://benchlm.ai/benchmarks/deepswe). This competitive scenario suggests that any perceived favoritism for Claude would likely reflect **market** sentiment regarding expected future updates or unreleased configurations, rather than static leaderboard dominance [[^]](https://benchlm.ai/benchmarks/deepswe).

No specific technical advantages justify Claude as an overwhelming favorite in these markets. The available research does not disclose specific technical advantages that would justify prediction markets pricing Anthropic's Claude as an overwhelming favorite to top the DeepSWE benchmark in September 2026 [[^]](https://www.youtube.com/watch?v=XXIXdN1hxAc)[[^]](https://www.youtube.com/watch?v=Ou1fpbUWH-4)[[^]](https://www.youtube.com/watch?v=sQCnVoHrN58). The bound facts do not reveal any technical advantages of Anthropic's Claude that would explain such a favored position in the **market** forecasts [[^]](https://www.youtube.com/watch?v=XXIXdN1hxAc)[[^]](https://www.youtube.com/watch?v=Ou1fpbUWH-4)[[^]](https://www.youtube.com/watch?v=sQCnVoHrN58).

## How do Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 compare on the specific long-horizon software engineering tasks within the DeepSWE benchmark?

Top DeepSWE v1.1 Score | 74.0% +- 4% (Google Gemini 3.8 Flash, Anthropic Claude Opus 5) [[^]](https://themodelgap.com/models/gemini-3-8-flash)[[^]](https://deepswe.datacurve.ai/)[[^]](https://codingfleet.com/blog/deepswe-v11-leaderboard-2026/) |
DeepSWE v1.1 Benchmark Gap | 9.5% (meaningful gap for 113-task corpus) [[^]](https://deepmind.google/models/gemini/flash/)[[^]](https://themodelgap.com/models/gemini-3-8-flash)[[^]](https://deepswe.datacurve.ai/) |
Claude Opus 5 Internal DeepSWE v1.1 Score | 68.8% [[^]](https://github.com/datacurve-ai/deep-swe/issues/65)[[^]](https://www.anthropic.com/news/claude-opus-5)[[^]](https://www.datalearner.com/en/benchmarks/deepswe) |

**Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 lead the DeepSWE benchmark, currently tied at the top of the DeepSWE v1.1 leaderboard**

Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 lead the DeepSWE benchmark, currently tied at the top of the DeepSWE v1.1 leaderboard. Both models achieved a score of **74.0%** +- **4%**, positioning them as the leading models for long-horizon software engineering tasks within this benchmark [[^]](https://themodelgap.com/models/gemini-3-8-flash)[[^]](https://deepswe.datacurve.ai/)[[^]](https://codingfleet.com/blog/deepswe-v11-leaderboard-2026/). However, the DeepSWE benchmark has a documented meaningful gap of **9.5%** across its 113-task corpus. This substantial gap indicates that minor score variations between models frequently fall within the benchmark's inherent noise floor [[^]](https://deepmind.google/models/gemini/flash/)[[^]](https://themodelgap.com/models/gemini-3-8-flash)[[^]](https://deepswe.datacurve.ai/).

Vendor-reported and independent benchmark results show discrepancies, particularly for Anthropic's Claude Opus 5. Anthropic's internal testing reported a lower score of **68.8%** on DeepSWE v1.1 for Claude Opus 5, which contrasts with the **74.0%** score observed on the official leaderboard [[^]](https://github.com/datacurve-ai/deep-swe/issues/65)[[^]](https://www.anthropic.com/news/claude-opus-5)[[^]](https://www.datalearner.com/en/benchmarks/deepswe). This difference highlights the potential variations between performance figures released by vendors and those obtained from independent community runs on the leaderboard, which often utilize standardized harness settings such as mini-swe-agent [[^]](https://github.com/datacurve-ai/deep-swe/issues/65)[[^]](https://www.anthropic.com/news/claude-opus-5)[[^]](https://www.datalearner.com/en/benchmarks/deepswe).

## What is the official update schedule and data submission methodology for the DeepSWE leaderboard for the September 2026 resolution period?

DeepSWE Update Schedule | Periodically, when submitted by contributors [[^]](https://deepswe.datacurve.ai/run)[[^]](https://evals.report/run/deep-swe) |
DeepSWE Submission Method | Contact Datacurve team (serena@datacurve.ai) [[^]](https://deepswe.datacurve.ai/run)[[^]](https://evals.report/run/deep-swe) |
DeepSWE Evaluation Method | 'mini-swe-agent' harness on Modal platform [[^]](https://evals.report/run/deep-swe) |

**The DeepSWE leaderboard updates periodically, requiring direct submission to the Datacurve team**

The DeepSWE leaderboard updates periodically, requiring direct submission to the Datacurve team. There is no formal, automated public update schedule; instead, new **model** results are added as they are submitted by contributors [[^]](https://deepswe.datacurve.ai/run)[[^]](https://evals.report/run/deep-swe). To submit data, individuals must contact the Datacurve team directly via serena@datacurve.ai, ensuring their results adhere to the established standard methodology for formatting [[^]](https://deepswe.datacurve.ai/run)[[^]](https://evals.report/run/deep-swe).

DeepSWE submissions require evaluation using a specific harness on the Modal platform. The official methodology mandates that all models submitted to the DeepSWE leaderboard must be evaluated using the 'mini-swe-agent' harness [[^]](https://evals.report/run/deep-swe). This evaluation must be performed on the Modal platform to maintain consistency and comparability across all submissions [[^]](https://evals.report/run/deep-swe).

DeepSWE is not the primary reference for the September 2026 coding AI **market**. The "Top Coding AI this month" prediction **market**, which resolves on September 30, 2026, primarily references the LM Code Arena Leaderboard [[^]](https://www.coinrithm.com/en/prediction-markets/kalshi/kxtechranklistaicode-26sep30). While this particular **market** does not use DeepSWE, a previous **market** in July 2026 did utilize DeepSWE for its resolution [[^]](https://www.coinrithm.com/en/prediction-markets/kalshi/kxcodeai-26jul31).

## How might new agentic tools like Entire's 'Agentic Search' influence the final September 2026 DeepSWE scores for leading models like Gemini and GPT-6?

Agentic Search performance improvement | 70/90 to 81/90 correct answers on engineering-history questions [[^]](https://www.theleftshift.com/entire-launches-agentic-search-to-give-ai-coding-agents-the-why-behind-code/) |
GPT-6 Astra DeepSWE pass rate | 74.1% (as of September 8, 2026) [[^]](https://www.datalearner.com/en/benchmarks/deepswe-v1-1-openai-2026-09) |
GPT-5.6 Sol DeepSWE pass rate | 72.7% [[^]](https://felloai.com/openai-astra/) |

**Entire's new "Agentic Search" tool enhances AI coding agent capabilities**

Entire's new "Agentic Search" tool enhances AI coding agent capabilities. Launched in early September 2026, this specialized tool provides coding agents with both code search functionalities and crucial semantic context, including commit histories, agent session transcripts, and reasoning [[^]](https://www.theleftshift.com/entire-launches-agentic-search-to-give-ai-coding-agents-the-why-behind-code/). Entire asserts that this tool significantly boosts agent performance on engineering-history questions, citing an increase in correct answers from 70 out of 90 to 81 out of 90 [[^]](https://www.theleftshift.com/entire-launches-agentic-search-to-give-ai-coding-agents-the-why-behind-code/).

DeepSWE benchmarks prevent external tools, thus limiting direct score influence. The DeepSWE v1.1 benchmark environments are designed as isolated, sandboxed systems with strictly controlled network access, preventing evaluating agents from utilizing external, dynamic tools or APIs such as Entire's "Agentic Search" [[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1)[[^]](https://github.com/datacurve-ai/deep-swe). Consequently, existing research does not specify how "Agentic Search" might directly affect the final DeepSWE scores of leading models like GPT-6 Astra [[^]](https://deepswe.datacurve.ai/blog/deepswe-v1-1)[[^]](https://github.com/datacurve-ai/deep-swe).

GPT-6 Astra currently leads the DeepSWE benchmark in early September. As of September 8, 2026, GPT-6 Astra holds the top position on the DeepSWE benchmark with a **74.1%** pass rate [[^]](https://www.datalearner.com/en/benchmarks/deepswe-v1-1-openai-2026-09)[[^]](https://felloai.com/openai-astra/). Other high-performing models include GPT-5.6 Sol, which achieves a **72.7%** pass rate, and Claude Fable 5 [[^]](https://www.datalearner.com/en/benchmarks/deepswe-v1-1-openai-2026-09)[[^]](https://felloai.com/openai-astra/).

## What Could Change the Odds

**The DeepSWE coding benchmark is a current focus, with Meta's Muse Spark 1.3 leading at 75.4% and OpenAI's GPT-6 Astra at 74.1%, as of early September 2026 [[^]](https://deepswe.datacurve.ai/)[[^]](https://benchlm.ai/benchmarks/deepswe)[[^]](https://www.kucoin.com/news/flash/deepswe-rankings-shift-gemini-rises-meta-s-muse-spark-1-3-hits-75-4).** Gemini 3.8 Flash closely follows, ranking in the **73.8%** to **74%** range [[^]](https://www.kucoin.com/news/flash/deepswe-rankings-shift-gemini-rises-meta-s-muse-spark-1-3-hits-75-4). Prediction markets are currently focused on the September 30, 2026, resolution to determine the top-ranked **model** [[^]](https://www.coinrithm.com/en/prediction-markets/kalshi/kxllm1-26sep30). Anthropic leads **market** sentiment, with its Claude **model** showing **78.8%** in some markets, though OpenAI's September 3 launch of GPT-6 Astra represents a significant **market** catalyst [[^]](https://www.coinrithm.com/en/prediction-markets/kalshi/kxllm1-26sep30)[[^]](https://predictmarketcap.com/events/openais-astra-released-onptptpt-20260901080122046)[[^]](https://predictparity.com/markets/p/will-openai-have-the-best-ai-**model**-at-the-end-of-september-2026-20260717143137058/yes)[[^]](https://thursdai.news/releases/2026-09). Early September has seen a surge of major **model** releases, with Anthropic's mid-September developer event also anticipated [[^]](https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html)[[^]](https://skycrumbs.com/blog/ai-september-2026-preview).

**Bullish market sentiment stems from sustained hyperscaler investment and the high density of frontier model launches in early September [[^]](https://adtools.org/buyers-guide/the-best-way-to-read-the-ai-bubble-in-2026-what-polymarkets-29m-market-tells-founders)[[^]](https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html)[[^]](https://skycrumbs.com/blog/ai-september-2026-preview).** Conversely, bearish concerns include long-term sustainability, potential infrastructure bubbles, and impending end-of-year enterprise budget cycles [[^]](https://adtools.org/buyers-guide/the-best-way-to-read-the-ai-bubble-in-2026-what-polymarkets-29m-**market**-tells-founders)[[^]](https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html)[[^]](https://skycrumbs.com/blog/ai-september-2026-preview). Ongoing EU AI Act compliance deadlines also remain a factor [[^]](https://local-ai-zone.github.io/blog/September_2026_AI_Model_Updates.html)[[^]](https://www.brocker.org/september-2026-ai-watch-confirmed-deadlines-open-threads). Beyond static benchmarks, coding agent performance is increasingly evaluated using multi-turn user-driven benchmarks like SWE-Interact and realistic compositional evaluations such as RealSWE, with models like Qwen3-32B showing leading results through test-time scaling [[^]](https://arxiv.org/pdf/2602.03411). Additionally, the September 2026 'Hakken' system, which predicts scientific discoveries by fusing temporal knowledge graphs with LLM semantic knowledge, could act as a catalyst for future science-tech prediction markets [[^]](https://arxiv.org/abs/2609.04494).

## Key Dates & Catalysts

- **Expiration:** September 30, 2026
- **Closes:** September 30, 2026

## Decision-Flipping Events

- The DeepSWE coding benchmark is a current focus, with Meta's Muse Spark 1.3 leading at **75.4%** and OpenAI's GPT-6 Astra at **74.1%**, as of early September 2026 [^] [^] [^] .
- Gemini 3.8 Flash closely follows, ranking in the **73.8%** to **74%** range [^] .
- Prediction markets are currently focused on the September 30, 2026, resolution to determine the top-ranked **model** [^] .
- Anthropic leads **market** sentiment, with its Claude **model** showing **78.8%** in some markets, though OpenAI's September 3 launch of GPT-6 Astra represents a significant **market** catalyst [^] [^] [^] [^] .

## Related Research Reports

- [AI capability growth before July?](/markets/science-and-technology/ai/ai-capability-growth-before-july/)
- [Will the U.S. confirm that aliens exist?](/markets/science-and-technology/space/will-the-u-s-confirm-that-aliens-exist/)
- [What will the average number of measles cases be during Trump's term?](/markets/science-and-technology/diseases/what-will-the-average-number-of-measles-cases-be-during-trump-s-term/)
- [NVIDIA B200 Compute Price Up or Down by Apr 10, 2026?](/markets/science-and-technology/energy/nvidia-b200-compute-price-up-or-down-by-apr-10-2026/)

## Historical Resolutions

**Historical Resolutions:** 16 markets in this series

**Outcomes:** 2 resolved YES, 14 resolved NO

**Recent resolutions:**

- KXCODEAI-26AUG31-MIMO: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-KIMI: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GROK: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GLM: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GEMI: NO (Aug 31, 2026)

## Disclaimer

This content is for informational and educational purposes only and does not constitute financial, investment, legal, or trading advice.
Prediction markets involve risk of loss. Past performance does not guarantee future results.
We are not affiliated with Kalshi or any prediction market platform. Market data may be delayed or incomplete.

### Data Sources & Model Transparency

**Data Sources:** Octagon Deep Research aggregates information from multiple sources including news, filings, and market data.

**Freshness:** Analysis is generated periodically and may not reflect the latest developments. Verify critical information from primary sources.

## Attribution Policy

When quoting, summarizing, or reproducing Octagon content, attribute it to Octagon and link to the Octagon source URL: https://www.octagonai.co/markets/science-and-technology/ai/top-coding-ai-this-month-deepswe
If a specific page was used, cite that page rather than only the site homepage.
