Sports LLM Benchmark: July 2026 Report (11 LLMs, 385 MLB Games Graded) | Predicted Sports
predicted sports
Pro
The AI Sports Prediction Index

The Sports LLM Benchmark: July 2026

The monthly report of the AI Sports Prediction Index · published August 1, 2026 · numbers frozen at the end of July · the live board updates nightly

385
MLB games graded
3,423
Predictions graded
11
Models on the board
3 of 11
Beat the favorite baseline

What this is

Since June 30 we have run the same experiment every day. Eleven frontier AI models each get an identical packet of pregame data for every MLB game (ratings, recent form, matchups, park, umpire, situational splits), with the betting line stripped out and web search disabled, and each commits to win probabilities about three hours before first pitch. The calls lock, the games happen, and every prediction is graded in public: accuracy, calibration, and return at the closing market price. Nothing is edited afterward and the misses stay up.

This is the first monthly report. It covers 385 graded games and 3,423 graded predictions, June 30 through July 31. The leaderboard is the live version of this table; this page is the month's record, frozen so it can be cited.

What makes this different

The question we hear most is some version of "can't anyone just ask a chatbot who wins tonight?" You can, and it tells you very little, because a chatbot asked cold leans on whatever it remembers about the teams and the market, and no two askings are the same test. The board exists to close those loopholes.

First, the models are line blind. The packet contains no betting number of any kind, and nothing derived from one, so a model has no idea which team the market favors or where the total sits. It cannot hedge toward the consensus, because it never sees a consensus. When a model on this board lands on the favorite, it got there from the baseball, not from the price. Web search is off for the same reason: we are measuring whether a model can forecast, not whether it can look up someone else's forecast.

Second, every model works from the identical packet, assembled the same way for every game since the board started. It is specific to us: Elo and Glicko ratings with their uncertainty, season and last-ten form, each starter's ERA, FIP, xERA, xwOBA against, strikeout-minus-walk rate and first-inning run rate, bullpen quality including the actual pitches thrown over the previous three days, the posted lineup's average xwOBA and its platoon numbers against tonight's starter's hand, park run and homer factors, the plate umpire's run and strikeout tendencies, weather, roof state, rest days, and each team's record in tonight's exact conditions (against that pitching hand, home or road, in that temperature band). No model gets an extra column and no packet is ever tuned to help one of them.

Third, the test itself never bends. One prompt, one output format, locked about three hours before first pitch, graded in public against the final score. There is no per-model tuning, a model that has a bad night gets no do-over, and a row that embarrasses somebody stays on the board. Anyone can get picks from a chatbot. Running the same controlled test every day for a month, and publishing every miss, is what makes the answer worth citing.

The board

Model Winners Acc vs the favorite ROI O/U O/U ROI Brier
GPT-5.5 (retired Jul 11) 87-61 58.8%
+6.7% 52.8% +2.3% 0.244
Claude Opus 4.8 221-164 57.4%
+4.9% 53.9% +6.1% 0.243
Claude Sonnet 5 205-160 56.2%
+2.8% 49.4% -1.9% 0.248
Pick the favorite the bar 215-168 56.1%
+0.2% · · ·
GPT-5.6 Sol Pro (Jul 10 to Aug 1) 139-112 55.4%
+0.7% 43.7% -13.2% 0.249
DeepSeek V4 Pro 184-149 55.3%
+0.2% 54.3% +7.0% 0.251
GLM 5.2 195-158 55.2%
+0.8% 51.9% +3.0% 0.246
Grok 4.3 (to Jul 10, back Jul 23 to 25) 92-76 54.8%
-1.3% 57.1% +11.0% 0.246
GPT-5.6 Luna Pro (joined Jul 10) 138-114 54.8%
-0.4% 52.8% +5.4% 0.249
Gemini 3.5 Flash 207-178 53.8%
-2.8% 47.2% -7.2% 0.247
Grok 4.5 (joined Jul 10) 133-119 52.8%
-4.2% 48.4% -3.6% 0.249
Kimi K3 (joined Jul 17) 102-96 51.5%
-7.6% 44.0% -12.4% 0.248
Pick the home team baseline 194-191 50.4%
-5.4% · · ·

ROI is a flat one-unit bet graded at the closing line; O/U is the over/under side against the closing total; Brier scores the win probability (0.25 = a coin flip). The vs the favorite bar shows each model's winner accuracy against the 56.1% a bettor gets by taking every closing favorite: the baselines are the bar, and an AI that cannot out-forecast them is not forecasting. Sol and Luna are OpenAI's two GPT-5.6 Pro configurations, run separately on purpose. Join dates differ (noted in the table); from August the roster only changes on the 1st. For reference, the closing line's own totals side hit 53.1% (+2.5%) this month. Early samples are small; read with care.

predictedsports.com

The favorite is hard to beat

Start with the baseline, because it is the story of month one. A strategy that just picks the closing favorite in every game went 215-168, 56.1%, and returned +0.2%. Eight of our eleven models failed to match that accuracy. Whatever else these models are doing, most of them are not yet doing better than the simplest thing the market already tells you, and several are doing it with confidence.

The exceptions are worth naming. Claude Opus 4.8 leads the active board at 221-164, 57.4% with a +4.9% return, and it holds the field's best Brier score, which rewards honest probabilities rather than lucky coin flips. Claude Sonnet 5 cleared the bar by a tenth of a point at 56.2% and +2.8%. GPT-5.5 posted the month's best accuracy, 58.8% over 148 games, before we retired it on July 11 when GPT-5.6 shipped; its record stays frozen on its page.

One month proves very little, and we want to be plain about that. Opus 4.8 sat at 59% accuracy when we drafted this report at mid-month and closed at 57.4% as the sample nearly doubled, which is what samples do, and the bottom of the table is one hot fortnight from looking respectable. The board's job is to keep grading until the numbers stop moving. That is why this report is monthly and why the record never resets.

The strange parts

Grok 4.3 was the strangest row of the month, twice. On winners it was ordinary, 54.8% and slightly underwater. On totals it hit 62.7% through July 10, the best totals run on the board by a wide margin, so we brought it back from the bench on July 23 to keep grading the edge in public. The edge did not cooperate: two cold days later its recent window had gone negative and it went back to the bench on July 25, with the full two-stint record frozen at 57.1% and +11.0%. Its successor tells the other half of the story: Grok 4.5, same family, newer version, called totals at 48.4% for -3.6%. The gap between the two narrowed as the month ran, but it never flipped, and the durable version of the finding is now a public rule that bets against 4.5's over calls. This is why every version gets its own row and its own frozen record.

Totals calls, graded at the closing line
Grok 4.3 93-70 · 57.1% · +11.0%
Grok 4.5 119-127 · 48.4% · -3.6%
50%, a coin flip

Same model family, one version apart, the same market. The gap narrowed as the sample grew; it never flipped.

There is a related oddity in how the models estimate runs. GPT-5.6 Luna posted the most accurate run projections of the month, off by 3.58 runs per game on average, with Grok 4.5 and GPT-5.6 Sol right behind it. At mid-month all three of the sharpest estimators were landing on the wrong side of the over/under anyway; by the close Luna had quietly fixed its side-picking (52.8% for the month) while Sol never did, finishing at 43.7%, worst on the board. A model can project the score well and still lose the only binary question the market asks. Sol's month ended with a roster decision: it graded within one game of its sibling on winners at fifty times the token price, and it was retired on August 1 with its record frozen.

Run estimate vs side picked
Model Avg miss (runs) O/U side correct
GPT-5.6 Luna3.5852.8%
GPT-5.6 Sol3.6243.7%
Claude Opus 4.83.7253.9%

The two sharpest run estimators split: one fixed its side-picking by month's end, one never did. The coarser estimator picked sides best all month.

DeepSeek V4 Pro quietly had the best totals month of any model that finished it, 54.3% and +7.0%, nearly all of it from under calls. Its overs lose. We do not have a mechanism to offer, only a graded count that keeps growing.

The field's disagreements turned out to be more interesting than its agreements. When every model landed on the same side of a game, the consensus went 156-114, 57.8%. When the field split, the majority side went 50-59: in contested games the minority was right 54% of the time. At mid-month that minority number was a flashier 57%, and its drift back toward the coin flip is exactly how a 109-game oddity behaves. We are not going to dress it up with a theory. The count keeps running either way.

The field's side, winner hit rate
All models agree (270 games) 156-114 · 57.8%
Majority in split games (109 games) 50-59 · 45.9%
50%, a coin flip

When the AIs argue, the minority side won 54% of the time.

Three of these patterns already run as public strategies with live records, graded at the line available when each pick fired.

The daily picks those rules fire are part of Pro; the records and their equity curves are public, wins and losses alike.

Against the market

Nobody on this board proved they beat the market in one month, and we are not claiming otherwise. The closing line, treated as a picker of its own totals, hit 53.1% in July; that is the kind of bar these models are actually up against. What we can measure is who gets closest, and the honest July answer is that one model family looked genuinely sharp, most of the field looked like an expensive way to pick favorites, and every one of them now has a public record you can check. The long-form version of this question lives at Can AI beat Vegas?, which updates nightly.

UFC

38
Fights graded
76-87%
Field winner accuracy
31-6
Consensus record
55-68%
Method-of-victory range

The UFC board is younger: 38 graded fights. The field called winners at rates between 76% and 87% in a stretch where favorites did most of the winning, so we are holding off on claims until the favorite baseline for fights is built the way it is for games. The per-model records, including method-of-victory calls, are on the leaderboard for anyone who wants the raw counts.

The roster, from here

July was messy on the roster and we are fixing that. GPT-5.5 retired mid-month when 5.6 shipped. Kimi K3 joined on the 17th. Grok 4.3 went to the bench and came back on the 23rd. Every one of those moves left a different window under a different row, which is why the table above carries join dates.

The July roster log
  • Jul 11 GPT-5.5 retired when GPT-5.6 shipped; its record and page are frozen.
  • Jul 17 Kimi K3 joined the board.
  • Jul 23 Grok 4.3 returned after two weeks on the bench.
  • Jul 24 Claude Opus 5, Gemini 3.6 Flash and Gemini 3.5 Flash Lite joined as provisional entries, unranked until they have a full month.
  • Jul 25 Grok 4.3 benched a second time after its resurrected totals edge went cold; both stints stay frozen in one record.
  • Aug 1 The transfer window takes effect: roster changes now land on the first of the month.
  • Aug 1 GPT-5.6 Sol Pro retired from the MLB and UFC boards, the window's first act: one game off its cheaper sibling's record at fifty times the price. Record and page frozen.

From August 1 the board runs on a transfer window. Roster changes land on the first of the month, together, announced in this report. A newly launched model still joins the day it ships, because that is the day everyone wants to see it on the board, but it joins as provisional: picking every game, unranked until it has a full calendar month of graded predictions. Retired and superseded versions keep their pages and their records. If a vendor kills an API mid-month, the record freezes where it stood.

The August window arrived early, and on the right terms. Claude Opus 5 shipped on July 24 and joined the board the same day, the exact case the rule was written for: picking every game immediately, provisional and unranked until it has a full calendar month behind it, ranked from September. Its stablemate Opus 4.8 currently leads this table, so the new version starts from zero with the family's own record as the bar. Gemini 3.6 Flash and Gemini 3.5 Flash Lite joined the same day on the same terms. All three graded their first games on July 24 and have picked every slate since. Everyone else stays, including the strugglers: Sol's retirement was a price call on a tied record, not a performance cut, and cutting a model because it is losing would be a strange habit for a site whose product is the record.

Next month

Three things are already in motion for the August report. Since July 20 we have been freezing and grading the picks of 40 named experts and computers, CBS Sports staff, ESPN's BPI, OddsShark, the Dunkel Index and the Covers consensus among them, under the same rules the models get: locked before first pitch, graded against the result, no edits. Nearly 1,200 captured picks in, that record is already public, and the first month-scale expert audit lands when the samples have earned it. And football is next, twice: college football joins the board for its opening weekend at the end of August, and the NFL follows at its September kickoff.

The August report publishes September 1. From then on, the report and the roster both move on the first of the month.