Which AI predicts sports best?
GLM 5.2, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.5, DeepSeek V4 Pro, Claude Sonnet 5, GPT-5.6 Sol Pro, and GPT-5.6 Luna Pro — head to head on real games. Each gets the identical line-blind data packet (no betting line, no web search), locks its call ~3 hours before first pitch, and is graded in public. No edits, no do-overs.
The honest finding so far: nobody beats the closing line consistently — not even the frontier models. The race worth watching is who gets closest.
MLB standings · ranked by Brier (lower = better)
178 graded games| # | Model | Record | AccAccuracy | Brier | ROI |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.8👑 | 105–73 | 59% | 0.241 | +8.6% |
| 2 | GLM 5.2 | 82–64 | 56% | 0.243 | +2.5% |
| 3 | Gemini 3.5 Flash | 100–78 | 56% | 0.246 | +2.3% |
| 4 | Claude Sonnet 5 | 93–65 | 59% | 0.247 | +9.3% |
| 5 | DeepSeek V4 Pro | 77–67 | 53% | 0.250 | -3.3% |
| 6 | GPT-5.6 Luna Pro | 24–21 | 53% | 0.255 | -0.3% |
| 7 | GPT-5.6 Sol Pro | 26–18 | 59% | 0.255 | +11.2% |
| 8 | Grok 4.5 | 23–22 | 51% | 0.255 | -4.7% |
| 9 | Pick the favoritebaseline | 103–73 | 59% | 0.415 | +4.4% |
| 10 | Pick the home teambaseline | 87–91 | 49% | 0.511 | -8.4% |
Brier = mean squared error of the win probability (0 = perfect, 0.25 = a coin flip). Accuracy = how often the model's side won. ROI = return on a 1-unit bet on each pick at the closing market price (so beating the vig means clearing ~0%). The baselines are the bar: an AI that can't out-forecast "pick the home team" isn't forecasting. Early samples are small — read with care.
Get our sports prediction daily newsletter
Featuring our models and analysis — the AI board, our own picks, and where they disagree. Free, in your inbox every morning.
Run totals · over/under, ranked by accuracy
| # | Model | O/U acc | Avg miss | ROI |
|---|---|---|---|---|
| 1 | DeepSeek V4 Pro | 56% | ±3.7 | +8.5% |
| 2 | Claude Opus 4.8 | 52% | ±3.7 | +1.0% |
| 3 | GLM 5.2 | 52% | ±3.8 | +2.7% |
| 4 | Claude Sonnet 5 | 48% | ±3.8 | -5.3% |
| 5 | Gemini 3.5 Flash | 47% | ±3.7 | -8.3% |
| 6 | GPT-5.6 Luna Pro | 32% | ±3.0 | -37.1% |
| 7 | Grok 4.5 | 30% | ±3.0 | -41.5% |
| 8 | GPT-5.6 Sol Pro | 21% | ±3.0 | -58.3% |
| Market (closing line)baseline | 55% | — | +4.6% |
Each model projects the game's total runs (line-blind); we grade its over/under call against the closing market line (a no-vig multi-book consensus via The Odds API where captured; Kalshi before that). O/U acc = how often that call was right. Avg miss = mean runs off the actual total. ROI = return on a 1-unit bet at the closing price — so break-even is the market's price (here the favored side runs ~59¢), not 50%: a model can top 50% and still lose to the vig. The Market (closing line) row is the bar to clear. Early samples are small — read with care.
🔒 See each model's projected total on every game — not just the scoreboard.
Start 7-day free trialFirst inning · run in the 1st (YRFI), ranked by Brier
| # | Model | YRFI acc | Brier |
|---|---|---|---|
| 1 | GPT-5.6 Luna Pro | 58% | 0.245 |
| 2 | Claude Sonnet 5 | 54% | 0.246 |
| 3 | Gemini 3.5 Flash | 52% | 0.246 |
| 4 | GPT-5.6 Sol Pro | 55% | 0.247 |
| 5 | Grok 4.5 | 58% | 0.247 |
| 6 | GLM 5.2 | 51% | 0.248 |
| 7 | DeepSeek V4 Pro | 54% | 0.249 |
| 8 | Claude Opus 4.8 | 49% | 0.251 |
| Always “no run”baseline | 50% | 0.497 |
Each model calls the probability of a run in the first inning (either team), line-blind. First innings are close to a coin weighted toward "no run", so the Always "no run" baseline is the bar — an AI only shows skill by beating it on Brier. It's also the one market where our own first-inning model claims a real (modest) edge, so this table is the fairest fight on the board. Early samples are small — read with care.
UFC — best fight forecasters
14 graded fightsSame idea, in the cage: every model gets the identical line-blind fight packet — no odds, no search — calls the winner and the method, and is graded on both. Follow the picks on every card.
| # | Model | Record | Winner acc | Method acc |
|---|---|---|---|---|
| 1 | Claude Opus 4.8 | 12–2 | 86% | 57% (14) |
| 2 | Grok 4.5 | 12–2 | 86% | 64% (14) |
| 3 | Claude Sonnet 5 | 13–1 | 93% | 79% (14) |
| 4 | Gemini 3.5 Flash | 13–1 | 93% | 57% (14) |
| 5 | GLM 5.2 | 11–3 | 79% | 64% (14) |
| 6 | GPT-5.6 Luna Pro | 11–3 | 79% | 50% (14) |
| 7 | DeepSeek V4 Pro | 12–2 | 86% | 71% (14) |
| 8 | GPT-5.6 Sol Pro | 11–3 | 79% | 57% (14) |
Winner acc = how often the model's pick won. Method acc = how often it called the finish type right (KO/TKO · Submission · Decision), over fights that reached a clean result; the count in parentheses is that sample. Early samples are small.
How it works
Line-blind. No model ever sees the betting line. Each produces its own win probabilities and run projections from the data alone, so the board reads forecasting skill — not an echo of the market.
One packet, no search. Every model gets the same point-in-time data (ratings, Statcast, bullpens, park/umpire, situational splits) ~3h before first pitch. No web search, so it's the models we're measuring.
Graded in public, no do-overs. Each model makes one call per game and we live with it — against the outcome (accuracy, Brier) and against the closing price (ROI). Our own house models get the same treatment on the accuracy page.
Where the Index goes next
The standings change every night. Get the update.
Who's hot, who's slipping, where the AIs disagree — free in your inbox every morning.