Science and Technology Polymarket August 13, 2026 Crowd ahead of press
DeepSeek V4 Pro reportedly outperforms Grok 4.6 in first public test
Next Grok Model (4.6+): Text Arena Debut?
Polymarket prices this 1440+ at 99%. The market is more confident than the current reporting.
A report from 36Kr says DeepSeek founder Liang Wenfeng's new DeepSeek V4 Pro delivered explosive performance against Elon Musk's Grok 4.6 in its first public test, escalating the U.S.–China AI race just as xAI pushes Grok up the Arena.AI leaderboard. The specific question at issue is whether the next Grok model designated version 4.6 or higher can debut on that leaderboard with an overall text score of 1440 or higher by December 31, 2026. Meanwhile, xAI's Grok Imagine Image 2.0 tool launched last week to broad coverage, earning a second-place ranking on the image-generation arena. Traders put the odds of the 1440+ debut at 99%, a level that has shifted +25.1 pts over the past week even though the reporting centers on a competitor's claims rather than Grok's own text benchmark. The coverage is thin and largely sourced to a single Chinese outlet's account of an unpublished comparison.
Background
Elon Musk founded xAI in 2023 to compete with OpenAI, Google, and Anthropic, releasing the Grok family of large language models through his social platform X and via API. The Arena.AI leaderboard, maintained by LMSYS, ranks models on human-preference votes across text and image tasks; a score of 1440+ on the text overall board would place a model among the top tier currently occupied by the likes of GPT-5 and Claude 4. Grok's image-generation arm, Grok Imagine, reached number two on the image arena last week, according to multiple outlets covering the launch of version 2.0. DeepSeek, founded by Liang Wenfeng, has emerged as China's leading open-weight AI lab, with its models repeatedly matching or exceeding Western counterparts on benchmark scores at lower reported training cost. The market resolves on the calendar date after a qualifying Grok model — version 4.6 or higher — first appears on the text leaderboard.
The precedent
- DeepSeek's V3 model, released in December 2024, matched or exceeded top Western models on several benchmarks at a reported training cost under $6 million, establishing the lab as a serious competitor.
- The LMSYS Chatbot Arena has ranked models by human-preference votes since its launch in 2023, with the top text models typically scoring above 1400 on the Elo-based leaderboard.
Context compiled by Crowdtells from the public record — verify before relying on it.
What the coverage agrees on
- xAI launched Grok Imagine Image 2.0 with precise editing and text rendering features
- Grok Imagine reached second place on the Arena image-generation leaderboard
- DeepSeek V4 Pro was publicly tested against Grok 4.6, per 36Kr
- Grok Imagine 2.0's API remains unavailable
Where sources diverge
- Whether DeepSeek V4 Pro significantly outperformed Grok 4.6 — claimed by 36Kr but unconfirmed by other outlets or by published benchmark results
How outlets frame it
- 36Kr: Frames the DeepSeek V4 Pro test as a surprise offensive by Liang Wenfeng against Elon Musk, emphasizing competitive rivalry and explosive performance in a head-to-head framing the other outlets do not address.
- Pasquale Pillitteri: Highlights that Grok Imagine Image 2.0's API is still missing, a practical limitation the other launch-coverage outlets omit or downplay.
What to watch
The key date is December 31, 2026, the resolution deadline. Before then, watch for xAI to release a Grok model numbered 4.6 or higher and for it to appear on the Arena.AI text leaderboard; its score at noon ET the following day determines the outcome. DeepSeek V4 Pro's own benchmark numbers and any public comparison with Grok 4.6 could shift perceptions of whether xAI accelerates its release schedule.
The numbers behind this
Polymarket prices this 1440+ at 99%.
24h +4.8 pts 7d +25.1 pts
$51K traded · $26.6K in the last day · $10.1K resting liquidity · $23.8K open interest
Resolves on: This market will resolve to "Yes" if the next SpaceXAI Grok model added to the Arena.AI Leaderboard (https://arena.ai/leaderboard/text/overall-no-style-control) has at least the specified score at 12:00 PM ET on the calendar date following the date on which it first appears on the leaderboard. Otherwise, this market will resolve to "No". A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a…
Pricing Polymarket 99%
Sources
- Grok's image tools get a professional upgrade with templates, precise editing, and an Arena number two ranking thenextweb.com
- xAI Ships Grok Imagine Image 2.0 With Precise Editing and a Top Arena Ranking unite.ai
- Grok Imagine Image 2.0: xAI Takes Second in Arena, API Still Missing pasqualepillitteri.it
- Grok Imagine Image 2.0: xAI Launches Its New AI Image Generation Tool With Magic Wand, Text Rendering and latestly.com
- Liang Wenfeng Launches Surprise Offensive Against Elon Musk: DeepSeek V4 Pro vs Grok 4.6 Showcases Explosive Performance in First Public Test eu.36kr.com
Frequently asked questions
Next Grok Model (4.6+): Text Arena Debut?
Polymarket prices this 1440+ at 99%. The market is more confident than the current reporting.
What do the sources agree on?
xAI launched Grok Imagine Image 2.0 with precise editing and text rendering features Grok Imagine reached second place on the Arena image-generation leaderboard DeepSeek V4 Pro was publicly tested against Grok 4.6, per 36Kr Grok Imagine 2.0's API remains unavailable
Where do the sources disagree?
Whether DeepSeek V4 Pro significantly outperformed Grok 4.6 — claimed by 36Kr but unconfirmed by other outlets or by published benchmark results
When does this market resolve?
This market resolves on: This market will resolve to "Yes" if the next SpaceXAI Grok model added to the Arena.AI Leaderboard (https://arena.ai/leaderboard/text/overall-no-style-control) has at least the specified score at 12:00 PM ET on the calendar date following the date on which it first appears on the leaderboard. Otherwise, this market will resolve to "No". A qualifying SpaceXAI model must have “Grok” in its displayed model name and be designated as version 4.6 or higher, regardless of capitalization or surrounding prefixes, suffixes, dates, or descriptors. For example, grok-4.6-high, grok-4.7-thinking, grok-5, or similar would qualify. Models whose displayed name does not include “Grok,” or which retain a…
How are these odds set?
Prediction-market odds are prices set by people trading real money on the outcome, so the price reads as the crowd’s implied probability — not a guarantee or financial advice.
AI-written briefing grounded in 5 sources and the live market, edited by Samuel Jo. Odds are crowd probabilities, not advice — how this works.