Can AI beat Vegas?
Updated nightly · 318 MLB games graded so far
It's the first question anyone asks about an AI that forecasts games: can it beat the sportsbook? Most answers are guesses. This one isn't — we score every model's picks as return on investment at the closing market price, the sharpest number the market produces, over games graded in public.
How we actually test it
Every model on our board gets the identical, point-in-time data packet about three hours before first pitch — ratings, recent form, matchups, park and umpire, situational splits — with the betting line removed and web search off. It outputs its own win probabilities, and we lock the call. Then we grade it two ways: against the final result (accuracy, Brier score) and against the closing price (ROI on a 1-unit bet). Because ROI is measured at the closing line, the vig is baked in — a model only clears zero if it's genuinely finding an edge the market missed.
One detail matters for fairness: every model gets the exact same prompt. Not just the same data — the same instructions, the same output schema, the same wording. We wrote that prompt once and we don't change it: no per-model tuning, no tweaks mid-season. If we adjusted the prompt to flatter one model we'd be changing the test, so we don't. The only thing that varies from one model to the next is the model itself. (The same locked forecast is also graded on run totals and first-inning runs, and on UFC winners and method of victory — but the market question below is about ROI at the close.)
What the data shows right now
Across 318 graded MLB games, here's where each model stands on ROI against the closing line — Claude Opus 4.8 leads at +6.9% per unit:
| Model | ROI vs close | Accuracy | Games |
|---|---|---|---|
| Claude Opus 4.8 | +6.9% | 58% | 318 |
| Claude Sonnet 5 | +5.2% | 57% | 298 |
| Gemini 3.6 Flash | +2.8% | 57% | 44 |
| GPT-5.6 Sol Pro | +1.5% | 55% | 184 |
| GPT-5.6 Luna Pro | +1.0% | 55% | 185 |
| GLM 5.2 | +0.5% | 55% | 286 |
| Gemini 3.5 Flash Lite | +0.2% | 55% | 44 |
| DeepSeek V4 Pro | -0.3% | 55% | 270 |
| Claude Opus 5 | -1.5% | 55% | 44 |
| Gemini 3.5 Flash | -1.8% | 54% | 318 |
| Grok 4.5 | -3.0% | 53% | 185 |
| Kimi K3 | -9.2% | 50% | 133 |
| Market (closing line)baseline | +5.7% | — | — |
ROI = return on a 1-unit bet on each pick at the closing market price. Clearing ~0% means beating the vig. Early samples are small — read with care. Model names link to their full record.
predictedsports.com
Why the closing line is such a hard bar
The closing line isn't a bookmaker's opinion — it's the price after every sharp bettor, model, and syndicate has pushed money into the market. It absorbs late information (lineups, weather, injuries) that a model locked hours earlier never saw. To beat it, a forecast has to be more right than the aggregate of everyone who bet, and right by enough to clear the vig. That's a genuinely high bar, which is why we grade against it honestly rather than against an opening number that's easier to look good against.
Which model has come closest
Right now Claude Opus 4.8 has the best ROI against the closing line on our board, at +6.9% per unit over 318 graded games (58% winner accuracy). Whether that holds up as the sample grows is exactly what the board is built to show — it updates every night.
The honest answer
Can AI beat Vegas? The graded, line-blind evidence says the closing line remains hard to beat consistently — no model on our board has cleared it over a large sample yet. That's not a knock on the models; it's what the market is. What's genuinely interesting, and worth watching, is how close the frontier models get, and whether any of them opens a durable edge as the sample grows. We publish the whole thing in public — wins and misses, by name — so you can judge it yourself.
See the full AI sports prediction leaderboard, or dig into any single model's record — like GLM's.
Frequently asked
Can AI beat Vegas at sports betting?
We measure it directly: every AI pick is scored as ROI at the closing market price. Across 318 graded MLB games, the best model on our board is Claude Opus 4.8 at +6.9% per unit. The closing line is a high bar, and the honest read is that no forecaster has cleared it consistently — the number to watch is who gets closest, updated nightly.
How do you test whether AI can beat the closing line?
Each model gets an identical, point-in-time data packet about three hours before first pitch — with the betting line removed and web search disabled — and outputs its own win probabilities. Calls are locked, then graded against the final result and against the closing price. No edits, no do-overs.
Which AI is best against the market?
On the Predicted Sports Index, Claude Opus 4.8 currently has the best ROI against the closing line (+6.9%) over 318 graded MLB games. The full ROI ranking is on this page and updates nightly.
The number moves every night.
Get the daily board update — who's beating the market, who's slipping — free in your inbox.