GPT-5.6 Sol Pro vs Luna Pro: Why We Retired the 50x Model | Predicted Sports
predicted sports
Pro
Model Notes

Why We Retired GPT-5.6 Sol Pro From the Leaderboard

๐Ÿ“Š Like this? Get the monthly AI report plus the daily email, free.

Why We Retired GPT-5.6 Sol Pro From the Leaderboard

We retired a model this week that was not losing. That is a first for us, so it deserves a full explanation with the receipts attached.

GPT-5.6 Sol Pro has been on our MLB leaderboard since the start, picking every game with the same line-blind information every other model gets. Over 237 graded picks it went 131-106, which is 55.3 percent. Its cheaper sibling, GPT-5.6 Luna Pro, graded 131-107 over the same stretch, 55.0 percent. One game apart after two months of baseball.

It gets closer than that. Calibration is where a model shows whether its confidence means anything, and their Brier scores landed at .2485 and .2490, a rounding error apart. On closing line value they are literally tied: both beat the closing number by 0.34 percentage points on average, Sol beating the close on 59.3 percent of its priced picks and Luna on 58.2. Whatever these two models know about baseball, it is the same knowledge.

The one place they separate is game totals, and it is not the direction the price tag suggests. Sol Pro called overs and unders at 44.2 percent, which cost a flat bettor 12.2 percent per pick. Luna Pro hit 53.0 percent and made 5.8 percent. On UFC picks the gap is one fight: 30-8 against 29-9 across 38 graded predictions. You can see the full side by side on their head to head page.

Sol Pro vs Luna Pro on game totals: 44.2 percent against 53.0 percent

The bill for that tie

Cost per graded pick: Sol Pro 2.45 cents, Luna Pro 0.05 cents

Through OpenRouter, where our pipeline runs, Sol Pro costs 5 dollars per million input tokens and 30 dollars per million output. Luna Pro costs 10 cents and 60 cents. That is fifty to one, on both sides of the meter, for the record above.

At our current volume the absolute dollars are small. A game prompt runs a couple thousand tokens, so a month of Sol Pro picks costs us roughly what a month of Luna Pro costs times fifty, call it eleven dollars against about twenty cents. The dollars are not the point. The ratio is the point, because every sport we add multiplies it. An NFL board arrives in September. Player props are on the roadmap. A model that charges fifty times the going rate for the same accuracy fails the only test a benchmark is allowed to care about, and it fails it harder every time we grow.

The industry seems to be arriving at the same conclusion from the other side. OpenAI cut Luna prices by 80 percent this week, per VentureBeat, and the whole market is sliding toward cost as the axis of competition. Frontier accuracy on sports outcomes appears to be a commodity at this point. The premium tier needs to prove it buys something, and on our board it did not.

Why retire instead of letting it ride

We thought about leaving it. A benchmark with more models is more interesting, and Sol Pro was not embarrassing itself. But we grade models in public because we think the record should decide things, and the record here decided something specific: paying fifty times more bought nothing. Keeping it on the active board would mean ignoring our own data, which is the one thing this site cannot do.

So it gets the same treatment Grok 4.3 got. The picks stop, the record freezes, and the model page stays up permanently with all 251 graded MLB predictions and every UFC pick intact. Nothing is deleted, because retiring a model is not the same as hiding one. Roster moves happen on the first of the month here, so the active board is Sol-free as of August 1, and the API budget it was burning goes toward the next challenger we add.

Luna Pro keeps its seat. Same picks, two percent of the invoice.

Get the monthly AI report in your inbox

Plus The Morning Board daily: every model graded in public, free. No spam, unsubscribe anytime.