AI Sports Prediction Leaderboard: Which AI Predicts Best? | Predicted Sports
predicted sports
Pro
The AI Sports Prediction Index

Which AI predicts sports best?

GLM 5.2, Claude Opus 4.8, Grok 4.5, DeepSeek V4 Pro, Claude Sonnet 5, GPT-5.6 Luna Pro, Gemini 3.6 Flash, and Gemini 3.7 Flash — head to head on real games. Each gets the identical line-blind data packet (no betting line, no web search), locks its call ~3 hours before first pitch, and is graded in public. No edits, no do-overs.

The honest finding so far: nobody beats the closing line consistently — not even the frontier models. The race worth watching is who gets closest.

New: the July 2026 benchmark report, the month's numbers frozen and citable.

🥇
Gemini 3.7 Flash
73–36
67% · Brier 0.226
🥈
GLM 5.2
237–171
58% · Brier 0.242
🥉
DeepSeek V4 Pro
220–152
59% · Brier 0.243

MLB standings · ranked by Brier (lower = better)

408 graded games · last 30 days
# Model Record Acc Brier ROI
1 Gemini 3.7 Flash👑 73–36 67% 0.226 +15.3%
2 GLM 5.2 237–171 58% 0.242 +2.0%
3 DeepSeek V4 Pro 220–152 59% 0.243 +4.4%
4 Grok 4.5 232–176 57% 0.243 -0.4%
5 Claude Sonnet 5 237–171 58% 0.244 +2.3%
6 Claude Opus 4.8 234–174 57% 0.244 +0.6%
7 Gemini 3.6 Flash 235–173 58% 0.244 +0.5%
8 GPT-5.6 Luna Pro 238–170 58% 0.245 +2.2%
9 Pick the favoritebaseline 242–166 59% 0.407 +3.5%
10 Pick the home teambaseline 227–181 56% 0.444 +3.7%

Brier = mean squared error of the win probability (0 = perfect, 0.25 = a coin flip). Accuracy = how often the model's side won. ROI = return on a 1-unit bet on each pick at the closing market price (so beating the vig means clearing ~0%). CLV = closing-line value: how far the market moved toward the model's side between the moment it picked and first pitch, in percentage points. Positive means it was on the right side of the move, which shows up long before win/loss does. A typical MLB line moves about 1.9 points, so these are fractions of the available move. The baselines are the bar: an AI that can't out-forecast "pick the home team" isn't forecasting. Early samples are small — read with care.

The raw rows behind this board are an open dataset (CC BY 4.0): every graded prediction with results, closing prices and CLV, refreshed weekly. Each month's numbers freeze into the monthly benchmark report.

Embed this board on your site

Copy-paste, free, updates itself. Add ?theme=dark or ?rows=5 to the iframe URL to taste.

predictedsports.com

Get our sports prediction daily newsletter

Featuring our models and analysis — the AI board, our own picks, and where they disagree. Free, in your inbox every morning.

Run totals · over/under, ranked by accuracy

# Model O/U acc Avg miss ROI
1 DeepSeek V4 Pro 50% ±3.5 +0.2%
2 GPT-5.6 Luna Pro 50% ±3.5 -0.7%
3 Gemini 3.6 Flash 49% ±3.5 -1.7%
4 Claude Opus 4.8 48% ±3.5 -3.5%
5 Grok 4.5 47% ±3.5 -6.0%
6 Claude Sonnet 5 47% ±3.5 -5.8%
7 GLM 5.2 47% ±3.5 -6.8%
8 Gemini 3.7 Flash 42% ±3.6 -15.7%
Market (closing line)baseline 47% -7.3%

Each model projects the game's total runs (line-blind); we grade its over/under call against the closing market line (a no-vig multi-book consensus via The Odds API where captured; Kalshi before that). O/U acc = how often that call was right. Avg miss = mean runs off the actual total. ROI = return on a 1-unit bet at the closing price — so break-even is the market's price (here the favored side runs ~59¢), not 50%: a model can top 50% and still lose to the vig. The Market (closing line) row is the bar to clear. Early samples are small — read with care.

predictedsports.com

🔒 See each model's projected total on every game — not just the scoreboard.

Start 7-day free trial
Cancel anytime · then $9/mo

UFC — best fight forecasters

63 graded fights

Same idea, in the cage: every model gets the identical line-blind fight packet — no odds, no search — calls the winner and the method, and is graded on both. Follow the picks on every card.

# Model Record Winner acc Method acc
1 Gemini 3.5 Flash 44–19 70% 41% (63)
2 Claude Sonnet 5 42–21 67% 38% (63)
3 GPT-5.6 Luna Pro 44–19 70% 46% (63)
4 Claude Opus 5 42–21 67% 44% (63)
5 GLM 5.2 43–20 68% 44% (63)
6 Pick the favoritebaseline 38–11 78%
7 Grok 4.5 39–24 62% 38% (63)
8 DeepSeek V4 Pro 39–24 62% 40% (63)
9 Gemini 3.7 Flash 6–7 46% 38% (13)

Winner acc = how often the model's pick won. Method acc = how often it called the finish type right (KO/TKO · Submission · Decision), over fights that reached a clean result; the count in parentheses is that sample. Pick the favorite is the baseline: take whoever the closing odds board made the favorite, every fight. A model that can't out-forecast it isn't forecasting. Its record covers only fights where we captured a closing no-vig price, so it runs on a smaller sample than the models — and because a naked side pick is scored as a 100%-confident call, it carries a worse Brier than its win rate suggests. Early samples are small.

predictedsports.com

Premier League — best soccer forecasters

9 graded matches

Line-blind soccer forecasting: models call 1X2 match winner (home, draw, away), total goals, and BTTS with live injury reports. Follow every fixture on the Premier League AI Pick Board.

# Model Record 1X2 acc BTTS acc O/U 2.5
1 DeepSeek V4 Pro 3–1 75% 25% (4) 75% (4)
2 Claude Sonnet 5 5–4 56% 33% (9) 67% (9)
3 GPT-5.6 Luna Pro 5–4 56% 33% (9) 67% (9)
4 Gemini 3.7 Flash 5–4 56% 33% (9) 67% (9)
5 Grok 4.6 5–4 56% 22% (9) 67% (9)
6 Claude Opus 5 3–6 33% 33% (9) 67% (9)
7 GLM 5.2 4–4 50% 38% (8) 75% (8)

How it works

Line-blind. No model ever sees the betting line. Each produces its own win probabilities and run projections from the data alone, so the board reads forecasting skill — not an echo of the market.

One packet, no search. Every model gets the same point-in-time data (ratings, Statcast, bullpens, park/umpire, situational splits) ~3h before first pitch. No web search, so it's the models we're measuring.

Graded in public, no do-overs. Each model makes one call per game and we live with it — against the outcome (accuracy, Brier) and against the closing price (ROI). Our own house models get the same treatment on the accuracy page.

Where the Index goes next

Live now: nightly MLB board, UFC fight board, run totals graded against the closing line, ROI at closing prices.
World Cup: our panel is calling every knockout match — follow the bracket.
The daily board update — who's hot, who's slipping, where the models disagree, every morning in the newsletter.
NFL for the 2026 season — the same benchmark, on the biggest board in sports.
More markets — first-five and prop forecasts, graded the same honest way.
The citable Index — embeddable standings and a monthly report on the state of AI sports forecasting.
The daily board update

The standings change every night. Get the update.

Who's hot, who's slipping, where the AIs disagree — free in your inbox every morning.