The story everyone repeated after November 2024 was simple: the markets saw it coming and the forecasters didn't. It made a good headline. It is also only half true, and the half that gets left out is the half that matters for a midterm.
Here is the part the headline gets right. Rajiv Sethi, an economist at Columbia, compared Polymarket's 2024 prices against the three big statistical models — The Economist, Silver Bulletin, and FiveThirtyEight — and found Polymarket was the best forecast for the presidential contest. Traders were pricing a Trump win as clearly more likely while the models sat far closer to a coin flip. On the top-line question, the money was less hedged and it was right.
Now the part that gets left out. In that same comparison, the models were better than the market on the popular vote, and the market did particularly poorly down the ballot — worst of all on Congressional races. Among the models themselves the spread was narrow, with FiveThirtyEight coming out ahead and Silver Bulletin behind, but none of that changes the shape of the result: one presidential contract is not evidence about thirty-four Senate races.
One good call is not a track record
A separate study is even less flattering. Joshua Clinton and TzuFeng Huang at Vanderbilt went through more than 2,500 markets and over $2 billion in transactions across the final five weeks of the campaign. Their accuracy numbers are not uniform — 93% of PredictIt markets beat chance, but that fell to 78% on Kalshi and 67% on Polymarket. They also found prices for identical contracts diverging across exchanges, and arbitrage opportunities peaking in the last two weeks before Election Day, when attention and volume were at their highest.
That last detail is the one worth sitting with. A market that is efficiently aggregating information does not leave the same contract at two different prices on two venues during its busiest fortnight. The authors' conclusion is blunt: the findings "challenge the view that prediction markets necessarily efficiently and accurately aggregate information about political outcomes."
So the honest summary of 2024 is not "markets won." It is: the market was the single best read on the presidency, the models were better on the popular vote and on Congress, and the market's pricing was messier than the efficient-aggregation story suggests. One cycle, one presidential contract. That is a sample size of one on the question people actually cite it for.
Why that makes a 2026 scoreboard worth building
Read those two findings together and you get an uncomfortable result for anyone selling certainty in either direction.
The place prediction markets looked weakest in 2024 — down-ballot Congressional races — is precisely what the 2026 midterms are. Thirty-four Senate seats and the whole House. If you were going to pick the worst possible evidence for "just trust the market" in a midterm, it would be the 2024 down-ballot record.
And yet the models have their own problem: a rating like "Leans Democratic" is a band, not a number, and it can sit unchanged for months while a live contract reprices every day on real money. Comparing them at all requires converting the band, which means the comparison is only as honest as the conversion you publish.
We are not going to resolve that argument with an opinion. So we are keeping score instead.
The board
Our Markets vs Pundits board puts the live Kalshi contract price for each 2026 Senate race next to the most recent rating a named forecaster published for that same seat, converts the rating through a ladder we publish on the page, and shows the gap. Every row carries the date the rating was made and a link to the source, because a scoreboard that attributes a call to a named outlet is worth exactly as much as the proof that they made it.
A few things it deliberately does not do:
- It does not take a side. A gap means the market and the forecaster disagree. It does not mean either is wrong. The election settles that, not us.
- It carries no position size and no recommendation. There is no edge claim anywhere on it. It is a comparison of two published numbers.
- It leaves cells blank rather than guessing. Kalshi lists most of these races by candidate name instead of by party. For an open seat with two challengers, we hold no record of which party each candidate belongs to — so those rows show no market number. Inferring a party from a name would eventually put the wrong party next to a real person on a public page. A missing number is recoverable. That is not.
The gaps are already interesting. As of today the widest is New Hampshire, where the contract prices the Democratic side near 90% while the most recent published rating still reads "Leans Democratic" — which, on the ladder, is 60%. Minnesota and Georgia show the same shape. That is not a trade idea; it is a thirty-point disagreement between two serious sources about the same seat, and in November one of them will have been closer.
How to read a gap without fooling yourself
The ladder is coarse on purpose. Ratings outlets publish bands, so converting one into a probability means picking a number they never said. A gap under roughly ten points is mostly an artifact of that coarseness rather than a real disagreement — the honest read of a 4-point gap is "these two basically agree."
The wide ones are different. A thirty-point gap is not a rounding artifact. It usually means one of two things: the rating has not been refreshed since something changed, or the contract has moved on thin volume and is ahead of itself. Both are worth knowing. Neither tells you what to do about it.
Come back in November and the board will have settled its own argument. That is the point of writing the numbers down in advance.
Market prices referenced above are live and move continuously — read the current number on the board, not this sentence. Ratings are the published work of their respective outlets, linked on every row of the board and used here for comparison with attribution.