What will be the top AI model this week?
Short Answer
1. Market Behavior & Drivers
- The "Great Model Rush" defines current intense AI competition.
- Claude Opus 4.6 immediately set new benchmarks, featuring huge context windows.
- OpenAI launched GPT-5.3-Codex-Spark, strategically diversifying hardware from NVIDIA.
- LMArena Elo ratings reveal significant shifts in the AI competitive landscape.
- Claude Opus 4.6 holds an availability advantage over preview-status Gemini 3 Pro.
- Leaderboard updates on February 13-14 will be key market catalysts.
Current Context
2. Price Chart
Historical Price (Probability)
3. Significant Price Movements
Notable price changes detected in the chart, along with research into what caused each movement.
📉 February 12, 2026: 17.0pp drop
Price decreased from 23.0% to 6.0%
Outcome: claude-opus-4-6
📉 February 10, 2026: 9.0pp drop
Price decreased from 20.0% to 11.0%
Outcome: claude-opus-4-6
📉 February 09, 2026: 74.0pp drop
Price decreased from 89.0% to 15.0%
Outcome: claude-opus-4-6
4. Market Data
Contract Snapshot
This market resolves to YES if a specific AI model is determined to be the "top AI model this week," and to NO if no such model is identified. The market pertains to the current week, with the year 2026 also mentioned. Specific criteria for determining the "top AI model" and any special settlement conditions are not detailed in the provided content.
Market Discussion
People are actively discussing and debating the "top AI model this week" amidst a crowded field of new releases and specialized advancements [1]. Prediction markets currently show strong favor for Anthropic's claude-opus-4-6-thinking as the top-ranked AI model for the week ending February 14, 2026 [2]. This comes during an unprecedented "Model Rush" in February 2026, with major launches including Google's Gemini 3 Pro GA, OpenAI's GPT-5.3, xAI's Grok 4.20, and various Chinese models like Qwen 3.5, creating intense competition and pushing AI capabilities in areas like agentic planning, real-time awareness, and specialized coding [3]. Beyond specific models, the debate extends to the efficacy of large, general-purpose models versus smaller, specialized AI tools, as well as the societal impact of AI, particularly concerning job displacement and ethical considerations [4]. Some discussions also anticipate future innovations beyond current Large Language Models (LLMs), suggesting they are not the final form of AI technology [5].
5. Which AI Models Lead Preliminary Elo Ratings in 2026?
| Claude Opus 4.6 Elo Rating | ~1490–1503 [1] |
|---|---|
| Gemini 3 Pro GA Elo Rating | ~1486–1492 [1] |
| GPT-5.2 Elo Rating (Incumbent) | ~1465–1473 [1] |
6. What Are the Key Adoption Barriers for Gemini 3 Pro and Claude Opus 4.6?
| Gemini 3 Pro Status | Preview [1] |
|---|---|
| Claude Opus 4.6 Status | General Availability (GA) on February 5, 2026 [1] |
| Gemini 3 Pro Base Token Cost | Approximately 60% lower than Claude Opus 4.6 [2] |
7. What Critical Reasoning Failures Plague Claude Opus 4.6 and Gemini 3 Pro?
| Claude Opus 4.6 Sabotage Hiding Success Rate | 18% [1] |
|---|---|
| Claude Opus 4.6 Injection Attack Success Rate | 50% [2] |
| Gemini 3 Deep Think ARC-AGI-2 Score | 45.1% [3] |
8. How Do Qwen 3.5 and GLM-5 Reshape the Open-Source LLM Landscape?
| GLM-5 Open-Source Date | February 11, 2026 [1] |
|---|---|
| GLM-5 Parameters | 744 billion total / 40 billion active (MoE) [1] |
| GLM-5 Leaderboard Rank | #1 among open-weight models (Artificial Analysis) [2] |
9. Does LMSys Chatbot Arena Have a Data Cutoff for Markets?
| Official Data Cutoff | Not officially defined by LMSys; platform operates continuously [1] |
|---|---|
| Leaderboard Update Frequency | Dynamic, near real-time or daily intervals [2] |
| Feb 13, 2026 Votes | Fully incorporated into Elo ratings before Feb 14 market resolution [2] |
10. What Could Change the Odds
Key Catalysts
Key Dates & Catalysts
- Expiration: February 14, 2026
- Closes: February 14, 2026
12. Historical Resolutions
Historical Resolutions: 50 markets in this series
Outcomes: 4 resolved YES, 46 resolved NO
Recent resolutions:
- KXTOPMODEL-26FEB07-QWEN3: NO (Feb 07, 2026)
- KXTOPMODEL-26FEB07-MIST: NO (Feb 07, 2026)
- KXTOPMODEL-26FEB07-GROK: NO (Feb 07, 2026)
- KXTOPMODEL-26FEB07-GPT5: NO (Feb 07, 2026)
- KXTOPMODEL-26FEB07-GPT: NO (Feb 07, 2026)