The Grok Fade: How One AI's Bad Overs Became Our Best Betting Strategy
There's a strategy on our Pro board right now called The Grok Fade, and the honest description of it is this: every time Grok 4.5 projects an MLB game to go over the total, we bet the under. No other conditions, no judgment calls, nothing clever layered on top.
Why would anyone do that? Well, because we grade every AI model on every game, in public, and the grades said something too loud to ignore. Grok 4.5's over calls, the games where it projected more runs than the betting line, hit 37% of the time. Twenty-six wins, fifty-two losses over its first month on our leaderboard. Flip a coin and you beat it comfortably. So we flipped the bet instead.

Where this came from
We run a weekly sweep over the graded record of every model we track. Not vibes, a literal backtest over every prediction and every closing number. The July 20 sweep flagged something in the totals data: a small cluster of models whose over calls were so bad they worked better inverted, and Grok 4.5 was the worst of them.
Here's the part that matters if you care about whether these things are real. We flagged the pattern on July 20, and THEN watched what happened next. On the 38 graded bets after that sweep, the fade returned +27%. That gap is the difference between finding a pattern and finding an edge, because a month of baseball is full of patterns and very few of them keep paying after you write them down.
By the July 26 sweep the full record was 52-26. As a bet, +34% return per pick at the pick-time line, across 78 games. We put it on the Pro board that day with the record frozen in public.
The checks we ran before trusting it
A month of unders winning is not automatically a Grok story, because July was an under month everywhere. Bet every under blindly in July and you made about 10% per pick. So the obvious question is whether the Grok Fade is just that tide wearing a costume.
It isn't, and the cleanest proof is the flat part of the market. In the 8 to 8.5 run band, where blind unders made basically nothing all month, games where Grok projected the over still cashed the under at a +35% clip. Games where Grok did NOT project the over? Betting the under there lost 19%. Same band, same month, opposite results, and the only difference is what one model said. A tide doesn't behave like that.

There's a mechanism under it too. Grok 4.5 is the only model on our board with negative closing line value, meaning the market moves AGAINST its side of the number more often than not (it lands on the right side of the move just 41.5% of the time). Its overall run projections aren't biased high, which is the strange part. So this isn't a model that leans toward offense in general. It is selectively wrong, in one direction, at the exact moments the market is most tempted to agree with it.
And this isn't even the first Grok totals story. Grok 4.3, the previous version, had a genuinely good totals edge, +21.7% on its picks. The 4.5 upgrade didn't just lose the edge, it inverted it. Model edges belong to exact versions. We learned that one in public too.
The honest part
Since we froze the record on July 26, the fade has gone 2-7. Nine bets means nothing statistically, and the strategy's live health monitor still grades it healthy on the trailing two weeks, but we publish the cold streaks with the same font size as the hot ones. As of this writing the live record is 54-33, +24.9% per pick, and you can watch it move every day on the Pro board, where every strategy carries its equity curve and a decay monitor that will flag this edge the moment it stops working.
The bigger idea is the one we keep coming back to. Everyone benchmarks AI models by asking what they get right. A graded record lets you profit from what a model reliably gets wrong, and reliably wrong is almost as valuable as reliably right. You just have to be willing to check the grades.
Every prediction behind this post is downloadable in our open dataset, and the methodology lives on the accuracy page. Check our math. That's what it's there for.

Plus The Morning Board daily: every model graded in public, free. No spam, unsubscribe anytime.