Top Coding AI this month (DeepSWE)
Short Answer
1. Market Behavior & Drivers
- ChatGPT leads; GPT-6 Astra scored 74.1% on DeepSWE in early September.
- Claude and Gemini remain strong from prior DeepSWE v1.1 leaderboard performance.
- New 'Agentic Search' tool may influence final DeepSWE scores for leading models.
Who Wins and Why
| Outcome | Market | Model | Why |
|---|---|---|---|
| Gemini | 29.0% | 29.2% | Gemini shows strong potential, integrating deeply with Google's ecosystem and tools. |
| ChatGPT | 56.0% | 40.8% | ChatGPT leads in coding capabilities, benefiting from frequent updates and continuous user feedback. |
| Claude | 35.0% | 23.0% | Claude is rapidly improving with its long context window, enhancing suitability for complex coding tasks. |
| Grok | 1.0% | 1.0% | Grok offers real-time information access, but its coding prowess remains unproven against leaders. |
| GLM | 2.0% | 2.0% | GLM models are robust for specific programming languages, yet lack broad adoption. |
Current Context
2. Price Chart
Historical Price (Probability)
3. Significant Price Movements
Notable price changes detected in the chart, along with research into what caused each movement.
Outcome: Claude
📉 September 08, 2026: 59.0pp drop
Price decreased from 94.0% to 35.0%
📈 September 07, 2026: 66.0pp spike
Price increased from 28.0% to 94.0%
📈 September 03, 2026: 12.0pp spike
Price increased from 14.0% to 26.0%
📉 September 02, 2026: 36.0pp drop
Price decreased from 50.0% to 14.0%
Outcome: ChatGPT
📈 September 01, 2026: 51.0pp spike
Price increased from 19.0% to 70.0%
4. Market Data
Contract Snapshot
This prediction market asks which AI model will be deemed the 'Top Coding AI this month (DeepSWE)'. The exact criteria for determining the 'Top Coding AI', which would trigger a YES resolution for that AI and a NO resolution for others, are not specified in the provided content. Trading begins on September 30, 12:00 AM EDT, and the market has a maximum payout date of September 30, 2026; no special settlement conditions are detailed.
Available Contracts
Market options and current pricing
| Outcome bucket | Yes (price) | No (price) | Last trade probability |
|---|---|---|---|
| ChatGPT | $0.75 | $0.74 | 56% |
| Claude | $0.80 | $0.95 | 35% |
| Gemini | $0.58 | $0.89 | 29% |
| GLM | $0.82 | $1.00 | 2% |
| Kimi | $0.95 | $1.00 | 2% |
| DeepSeek | $0.94 | $1.00 | 1% |
| Grok | $0.97 | $1.00 | 1% |
| MiMo | $0.82 | $1.00 | 1% |
Market Discussion
As of early September 2026, Meta's Muse Spark 1.3 is reported to lead the DeepSWE benchmark at 75.4% pass@1, with Gemini 3.8 Flash (73.8%) and GPT-6 Astra (74.1%) also featuring prominently in recent snapshots [^][^][^]. However, prediction markets strongly favor Anthropic's Claude models to secure the #1 position on other major coding benchmarks like LiveBench and Code Arena by the end of September 2026 [^][^][^][^]. Developer discussions in September 2026 indicate a shift from simple code generation metrics to focus on the reliability, verification, and cost economics of autonomous coding agents, with Meta's Muse Spark 1.3 specifically noted for its efficiency improvements [^][^][^][^][^][^][^].
5. Trader Dashboard
A deterministic, per-market integrity scorecard computed from order-book and price data. Higher is better for Trader Trust, Liquidity, Move Quality and Resolution; higher means more risk for Quote Risk and Avoid Risk.
GeminiPrimaryTrader Trust—Liquidity—Move Quality77Resolution—Quote Risk—Avoid Risk—
- Factor
- Factor
- flow_agreement
- 0.85
- move_retained_pct
- 100
ClaudeTrader Trust—Liquidity—Move Quality73Resolution—Quote Risk—Avoid Risk—
- Factor
- Factor
- flow_agreement
- 0.85
- move_retained_pct
- 100
ChatGPTTrader Trust—Liquidity—Move Quality60Resolution—Quote Risk—Avoid Risk—
- Factor
- Factor
- flow_agreement
- 0.85
- move_retained_pct
- 83
GrokTrader Trust—Liquidity—Move Quality79Resolution—Quote Risk—Avoid Risk—
- Factor
- Factor
- flow_agreement
- 0.85
- move_retained_pct
- 100
DeepSeekTrader Trust—Liquidity—Move Quality—Resolution—Quote Risk—Avoid Risk—
- Factor
- Metric
- —
KimiTrader Trust—Liquidity—Move Quality—Resolution—Quote Risk—Avoid Risk—
- Factor
- Metric
- —
MiMoTrader Trust—Liquidity—Move Quality—Resolution—Quote Risk—Avoid Risk—
- Factor
- Metric
- —
GLMTrader Trust—Liquidity—Move Quality—Resolution—Quote Risk—Avoid Risk—
- Factor
- Metric
- —
trader_dashboard_lean_v1.14 · computed Sep 8, 2026
6. What upcoming model releases from Anthropic, Google, or OpenAI could alter the DeepSWE leaderboard before September 30, 2026?
| GPT-6 Astra DeepSWE Score | 74.1% [^] |
|---|---|
| DeepSWE Leaderboard Update | September 3, 2026 [^] |
| DeepSWE Prediction Market Target | 90% Pass@1 [^][^] |
7. What technical advantages justify prediction markets pricing Anthropic's Claude as the overwhelming favorite to top the DeepSWE benchmark in September 2026?
| Claude Opus 5 DeepSWE v1.1 Score | 68.8% [^][^] |
|---|---|
| Sep 2026 DeepSWE Prediction - Gemini 3.8 Flash | 73.8% [^] |
| Sep 2026 DeepSWE Prediction - Claude Opus 5 | 73.6% [^] |
8. How do Google's Gemini 3.8 Flash and Anthropic's Claude Opus 5 compare on the specific long-horizon software engineering tasks within the DeepSWE benchmark?
| Top DeepSWE v1.1 Score | 74.0% +- 4% (Google Gemini 3.8 Flash, Anthropic Claude Opus 5) [^][^][^] |
|---|---|
| DeepSWE v1.1 Benchmark Gap | 9.5% (meaningful gap for 113-task corpus) [^][^][^] |
| Claude Opus 5 Internal DeepSWE v1.1 Score | 68.8% [^][^][^] |
9. What is the official update schedule and data submission methodology for the DeepSWE leaderboard for the September 2026 resolution period?
| DeepSWE Update Schedule | Periodically, when submitted by contributors [^][^] |
|---|---|
| DeepSWE Submission Method | Contact Datacurve team (serena@datacurve.ai) [^][^] |
| DeepSWE Evaluation Method | 'mini-swe-agent' harness on Modal platform [^] |
10. How might new agentic tools like Entire's 'Agentic Search' influence the final September 2026 DeepSWE scores for leading models like Gemini and GPT-6?
| Agentic Search performance improvement | 70/90 to 81/90 correct answers on engineering-history questions [^] |
|---|---|
| GPT-6 Astra DeepSWE pass rate | 74.1% (as of September 8, 2026) [^] |
| GPT-5.6 Sol DeepSWE pass rate | 72.7% [^] |
11. What Could Change the Odds
Key Catalysts
Key Dates & Catalysts
- Expiration: September 30, 2026
- Closes: September 30, 2026
12. Decision-Flipping Events
- Trigger: The DeepSWE coding benchmark is a current focus, with Meta's Muse Spark 1.3 leading at 75.4% and OpenAI's GPT-6 Astra at 74.1%, as of early September 2026 [^] [^] [^] .
- Trigger: Gemini 3.8 Flash closely follows, ranking in the 73.8% to 74% range [^] .
- Trigger: Prediction markets are currently focused on the September 30, 2026, resolution to determine the top-ranked model [^] .
- Trigger: Anthropic leads market sentiment, with its Claude model showing 78.8% in some markets, though OpenAI's September 3 launch of GPT-6 Astra represents a significant market catalyst [^] [^] [^] [^] .
14. Historical Resolutions
Historical Resolutions: 16 markets in this series
Outcomes: 2 resolved YES, 14 resolved NO
Recent resolutions:
- KXCODEAI-26AUG31-MIMO: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-KIMI: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GROK: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GLM: NO (Aug 31, 2026)
- KXCODEAI-26AUG31-GEMI: NO (Aug 31, 2026)