Claude vs ChatGPT Sports Predictions: Head-to-Head, Graded | Predicted Sports
predicted sports
Pro
Analysis

Claude vs ChatGPT at Sports Predictions: We Built the Page That Settles It

๐Ÿ“Š Like this? Get the monthly AI report plus the daily email, free.

"Which AI is smarter" arguments usually run on benchmarks nobody can check and vibes nobody can grade. We wanted the dumber, better version of that argument, two pickers on the same games with real finals and a count of who was right when they disagreed, so we built it. There are now 171 head-to-head pages live at predictedsports.com/vs, one for every pair of AI models and tracked experts on our board. Claude vs ChatGPT. Grok vs Gemini. Any AI against ESPN's BPI or the CBS Sports staff. Pick any two, and the page shows their record on the exact games both graded, updated nightly.

Opus 4.8 vs Grok 4.5: 130-121 across 209 shared games, 19-10 on disagreements

How the comparison actually works

The problem with most model comparisons is schedule bias. One model's record includes games another never picked, so the "records" aren't comparable. Our pages only count games BOTH pickers graded, so the schedule is identical by construction. Same slates, same finals, no cherry-picking possible, and we say so right on the page.

Then there's the number we care about most: the disagreement record. On games where two pickers took the SAME side, their results are identical by definition, so agreement tells you nothing about who's better. All of the information is in the games where they took opposite sides, and one of them had to be right. That's the headline stat on every page.

Each page also carries the windows (all time, last 30, last 7) and each picker's closing line value, which is the forward-looking stat: how far the market moved toward a pick after it was made. Accuracy tells you who was right, and CLV tells you who was early.

What the pages say right now

A few results from the current board, late July, all of it moving nightly:

Claude Opus 4.8 vs Grok 4.5 is the marquee AI matchup, and it's closer than the leaderboard suggests: 130-121 to Opus across 209 shared games. But on the 29 games where they took opposite sides, Opus leads 19-10. Nearly two out of three disagreements went Claude's way. That's the stat that separates them.

Claude Opus 4.8 vs GPT-5.6 Sol Pro (the current ChatGPT flagship on our board) sits 129-125 on 208 shared games, with Opus taking the disagreements 16-12. Tight, and genuinely live, which is exactly why the page exists.

Opus 4.8 vs GPT-5.6 Sol Pro: 129-125 on 208 shared games, disagreements 16-12

The humbling ones involve the humans. Opus 4.8 vs ESPN BPI is dead even, 34-34 on shared games, disagreements split 5-5. And GLM 5.2, a legitimately capable model, trails the CBS Sports staff 41-34 on their shared slate, with CBS winning the disagreements 13-6. Anyone telling you the AIs have simply lapped the expert industry has not graded the games. Ours are all on the expert records page, scraped from public picks before first pitch and never edited.

Why we built 171 of these

Partly because the data was already there and refusing to compute it felt like a waste. Every prediction on the leaderboard is timestamped, line-blind, and graded against the final, so any pairing is just a different cut of the same ledger.

But mostly because this is the argument people actually have. Nobody searches "what is Claude's Brier score." They search Claude vs ChatGPT, and every answer they'll find is a benchmark screenshot from a vendor deck. We'd rather the answer be a scoreboard that updates nightly, shows the losses, and covers fights the vendors would never stage, like their model against a newspaper's picks.

One honest caveat, stated on every page: a month of shared games is a young sample, and several of these matchups are inside the noise. The pages carry the sample sizes for exactly that reason. When a new model launches it starts a page against everyone on day one, provisional until it has a full month, and its first disagreements start counting immediately.

Pick your favorite matchup at /vs, or start from any model's own record page and follow the head-to-head links. If you think we graded something wrong, the raw predictions are in the open dataset, so you can settle it with receipts.

Get the monthly AI report in your inbox

Plus The Morning Board daily: every model graded in public, free. No spam, unsubscribe anytime.