Original research · 64,507 contracts · Polymarket & Kalshi
The longshot tax
Cheap prediction-market contracts win far less often than their price says they will. I measured it across 64,507 resolved contracts on two venues. The pattern is real, it is large, and at prices you can actually trade it costs you more than half your money.
A prediction-market contract is the simplest financial instrument there is. You pay some price between zero and one dollar. If the thing happens, you get exactly one dollar. If it doesn't, you get nothing. That's it.
Because the payout is fixed at a dollar, the price is the market's probability estimate. A contract trading at 8¢ is the market saying "about an 8% chance." This is why journalists quote Polymarket odds as though they were forecasts; the translation really is that direct.
So here's a question you can actually answer with data: when the market says 8%, how often does it happen?
I took every resolved contract I could collect from Polymarket and Kalshi, recorded what each one cost 24 hours before it resolved, then checked what actually happened. 64,507 contracts. The answer is that above roughly 10¢ these markets are impressively well calibrated. Below 10¢ they are not, and the cheaper the contract, the worse it gets.
Slightly over half the money, gone. Not on some exotic strategy, just the simple act of buying cheap contracts at the price the screen was showing.
#What favourite–longshot bias is, and why anyone cares
This pattern has a name and a long history. In 1988 Richard Thaler and William Ziemba published a paper in the Journal of Economic Perspectives documenting it in horse racing: heavy favourites were priced roughly fairly, while extreme longshots, horses at 100-to-1, returned something like sixty cents on the dollar. Bettors systematically overpaid for small chances.
It became one of the most replicated anomalies in economics, and the argument since has been about why. Broadly, two camps:
- ▸Psychology. People are bad at small probabilities. Almost nobody can really tell 1% from 3%, and both get overweighted. Snowberg and Wolfers made the strongest version of this case in 2010, using exotic bets to separate a genuine taste for risk from simple misperception. The misperception story fit better.
- ▸Market structure. The bias doesn't need irrational people at all. Hyun Song Shin showed in the early 1990s that a bookmaker facing better-informed bettors will shade longshot odds defensively, producing the same pattern from entirely rational participants.
The lottery in the room
None of this surprises anyone who has ever stood in line for a Powerball ticket. Nobody feels the two cents. A loss that small does not register as a loss at all, and set against it is a thin promise of winning big, which is the part actually being bought. Cheap contracts sell a story, and stories get priced by how good they feel, not by how often they come true. A 2¢ contract is a lottery ticket with an order book attached.
Prediction markets are a genuinely new place to test this. They're not bookmakers; they're order books where users trade with each other, and the venue takes a fee and is indifferent to who wins. If the bias is about bookmaker behaviour, it should be weaker here. If it's about human psychology, it should follow the humans.
#How I measured it
Three choices did most of the work, and each is a place where this analysis could have gone wrong.
I priced each contract 24 hours before it resolved
This sounds like a detail. It is the whole study.
The obvious approach, using each market's last traded price, produces a spectacular and completely fake result. I checked: contracts whose final trade printed at 1¢ went on to win 0.245% of the time, about a quarter of what 1¢ implies. That looks like enormous longshot bias. It isn't. The last trade in a market happens once the outcome is essentially known: someone dumping a doomed position minutes before settlement. It measures the market converging on an answer it already has, not the market forecasting.
So every price here is sampled at a fixed point before resolution: 24 hours out for the headline numbers, and I verified every single row was drawn at or before its cutoff. I also ran 1 hour, 7 days and 30 days as checks.
I measured the gap in relative terms, not percentage points
If a 1¢ contract wins 0.7% of the time instead of 1%, the gap is 0.3 percentage points, which looks like nothing next to a 45¢ contract missing by 5 points. But the 1¢ buyer lost 30% of their money and the 45¢ buyer lost 11%. Percentage points systematically hide tail effects. Everything below is relative shortfall: how far realized frequency fell short of the price, as a fraction of the price.
This isn't a stylistic choice. An earlier version of this analysis measured in percentage points and concluded Polymarket had no bias at all. That conclusion was an artefact of the units.
I accounted for contracts that come in groups
Many markets are one event with many outcomes: twenty candidates for a nomination, say. Those outcomes aren't independent: exactly one can win. Treating them as separate observations badly overstates how much data you have. Every statistic here clusters by event.
#Polymarket: 44,116 contracts
Well calibrated across most of the range, with a clear tail problem below 5¢.
Polymarket, priced 24 hours before resolution
| Price | Contracts | Winners | Priced at | Actually won | Shortfall |
|---|---|---|---|---|---|
| under 1¢ | 4,735 | 13 | 0.40% | 0.27% | −31.3% |
| 1–2¢ | 1,513 | 15 | 1.42% | 0.99% | −30.2% |
| 2–5¢ | 2,433 | 69 | 3.27% | 2.84% | −13.3% |
| 5–10¢ | 2,246 | 157 | 7.13% | 6.99% | −2.0% |
| 10–20¢ | 3,977 | 620 | 14.91% | 15.59% | +4.6% |
| 20–35¢ | 8,392 | 2,408 | 27.27% | 28.69% | +5.2% |
| 35–50¢ | 8,512 | 3,727 | 42.90% | 43.79% | +2.1% |
| 50–65¢ | 6,251 | 3,512 | 55.97% | 56.18% | +0.4% |
| 65–80¢ | 3,150 | 2,227 | 71.79% | 70.70% | −1.5% |
| 80–90¢ | 1,080 | 908 | 84.43% | 84.07% | −0.4% |
| 90–97¢ | 705 | 675 | 93.63% | 95.74% | +2.3% |
| 97¢+ | 1,122 | 1,119 | 99.03% | 99.73% | +0.7% |
40,783 distinct events. Read the middle of this table as a compliment: between 10¢ and 97¢ Polymarket is accurate to within a few percent of its own price. That is genuinely good forecasting.
The tail is a different story. The two cheapest buckets, 6,248 contracts between them, came in about 30% below what their prices implied.
#Kalshi: 20,391 contracts
Same shape, much larger, and it extends far further up the price range.
Kalshi, priced 24 hours before resolution
| Price | Contracts | Winners | Priced at | Actually won | Shortfall |
|---|---|---|---|---|---|
| under 1¢ | 2,781 | 2 | 0.49% | 0.07% | −85.3% |
| 1–2¢ | 2,019 | 10 | 1.14% | 0.50% | −56.6% |
| 2–5¢ | 2,214 | 32 | 2.95% | 1.45% | −51.0% |
| 5–10¢ | 1,581 | 75 | 7.04% | 4.74% | −32.6% |
| 10–20¢ | 1,641 | 174 | 14.32% | 10.60% | −26.0% |
| 20–35¢ | 1,854 | 426 | 26.82% | 22.98% | −14.3% |
| 35–50¢ | 1,732 | 610 | 43.02% | 35.22% | −18.1% |
| 50–65¢ | 1,383 | 863 | 56.97% | 62.40% | +9.5% |
| 65–80¢ | 1,240 | 904 | 72.58% | 72.90% | +0.4% |
| 80–90¢ | 1,090 | 960 | 84.77% | 88.07% | +3.9% |
| 90–97¢ | 1,231 | 1,177 | 93.63% | 95.61% | +2.1% |
| 97¢+ | 1,625 | 1,612 | 98.61% | 99.20% | +0.6% |
2,387 distinct events. Note the mirror image: everything below 50¢ underperforms its price, everything above 50¢ overperforms. That is the textbook favourite–longshot signature (longshots overpriced, favourites underpriced), and here it is on a regulated US exchange.
Kalshi shortfall by price bucket
Kalshi: how far realized frequency fell short of (red) or exceeded (green) the market price, as a share of the price. Centre line is perfect calibration.
#The two venues look different. Mostly they aren't.
Kalshi's bias looks several times larger than Polymarket's, and the tempting story is that something about Kalshi causes it. I spent a lot of effort on that story and it did not hold up.
The problem is that the two venues list different things. Kalshi's sample is dominated by daily financial and commodity settlements; Polymarket's by politics, sports and crypto. Comparing them directly compares subject matter as much as venue. On top of that, Kalshi contracts arrive in long ladders, twenty strikes on the same underlying, while 93% of my Polymarket sample is standalone markets. The two datasets are shaped differently in ways that leak into any naive comparison.
So I used a test that removes the problem by construction. It's from Mukhtar Ali's 1977 work on racetrack betting, and it only ever compares contracts within the same event: the twenty strikes on one underlying, competing for the same single winner. Everything that differs between events (venue rules, subject, date, tick size, liquidity) is held fixed automatically, because you're only ever comparing legs against their own siblings.
It produces one number, s. If s = 1, prices inside the event are perfectly calibrated against each other. If s > 1, cheap legs win less often than their prices imply relative to dear legs: favourite–longshot bias, cleanly measured.
Within-event bias, both venues
| Venue | Events | Contracts | s | Std. error | t |
|---|---|---|---|---|---|
| Polymarket | 1,099 | 2,323 | 1.396 | 0.105 | +3.76 |
| Kalshi | 1,368 | 9,801 | 1.250 | 0.043 | +5.86 |
| Kalshi, at the asking price | 1,138 | 8,495 | 1.362 | 0.048 | +7.59 |
Both significantly above 1. The difference between them is not statistically significant (t = 1.29, p = 0.20), and the point estimate is actually higher on Polymarket.
The finding
Once you compare like with like, the two venues have the same bias. The apparent gap was composition: what each venue happens to list, not how each venue works.
This also killed my leading explanation. I had expected the bias to track the venue's minimum price increment: Kalshi quotes in whole cents, so a genuine 0.3% event literally cannot be priced honestly there, while Polymarket can quote far finer. That predicts much more bias on Kalshi. I got the opposite ordering, and the effect survives deleting every contract under 5¢, the only region where a one-cent floor can bind at all. The tick explanation is dead.
What's left standing is the older, simpler explanation: people overpay for small probabilities, and they do it about equally wherever you let them.
#What it costs at prices you can actually trade
Everything above uses the market's quoted price. But you can't buy at the midpoint; you buy at the ask, and on cheap contracts the spread is punishing. Kalshi's median spread is 3¢; at the 90th percentile it's 24¢. On a contract nominally worth 5¢, that is not a detail.
So here is the same analysis using the actual asking price, what a buyer really paid.
Kalshi: buying at the ask, 24 hours out
| Ask price | Contracts | Winners | Mean ask | Actually won | Return per $1 | 95% CI |
|---|---|---|---|---|---|---|
| under 1¢ | 170 | 0 | 0.56% | 0.00% | −100% | wide |
| 1–2¢ | 3,497 | 5 | 1.01% | 0.14% | −87.0% | [−95%, −67%] |
| 2–5¢ | 2,282 | 17 | 2.88% | 0.74% | −76.6% | [−85%, −59%] |
| 5–10¢ | 1,791 | 57 | 6.67% | 3.18% | −51.9% | [−64%, −39%] |
| 10–20¢ | 1,847 | 144 | 14.10% | 7.80% | −45.7% | [−53%, −35%] |
| 20–40¢ | 2,363 | 514 | 29.03% | 21.75% | −25.9% | [−31%, −19%] |
| 40–60¢ | 1,657 | 648 | 49.03% | 39.11% | −20.9% | [−25%, −15%] |
Every bucket is negative, and the confidence intervals exclude zero everywhere below 40¢. These are before Kalshi's trading fees, which make it slightly worse.
Two things stand out. First, the losses are enormous: a 1–2¢ contract bought at the ask returned 13 cents on the dollar. Second, the bias extends much further up the price range than the midpoint analysis suggested: even 40–60¢ contracts, the coin-flips, lost 21% when bought at the ask. That's the spread eating you on top of the mispricing.
#Does it hold up?
Three checks I'd want to see if someone showed me this.
Does it depend on when you measure?
No. Measuring at 1 hour, 24 hours, 7 days and 30 days before resolution, the sub-10¢ shortfall on Kalshi runs −44%, −45% and −34%. On Polymarket, −14%, −9%, −10% and −50%. Same sign every time on both venues.
Is it one weird corner of the market?
No. On Kalshi, excluding the auto-generated multi-leg parlay contracts changes the result not at all, and dropping any one of the eight largest contract families leaves it intact. It isn't one runaway category.
Could it be an artefact of how you built the data?
This is the one that nearly got me. Three independent reviewers went after this analysis specifically to break it, and they succeeded against an earlier version: an earlier headline claimed the bias appeared on Kalshi and not Polymarket, and that turned out to rest on comparing two different time horizons, two different liquidity filters, and two differently-constructed price series. It was withdrawn. The within-event result above is what survived, and it survived because it doesn't depend on any of those choices.
Where I'd push back on myself
The Polymarket within-event estimate rests on 1,099 events, because most Polymarket markets are standalone rather than multi-outcome. It is significant, but it's the thinnest number here; the claim it supports is "statistically indistinguishable from Kalshi," not "measurably larger."
And 67% of Kalshi's quoted prices in my data are midpoints of a quote rather than actual trades. That's why the executable-ask table matters more than the midpoint table: the ask is a real price someone would have filled you at.
#What this means if you trade these markets
- 01Above roughly 10¢, treat the price as a decent forecast. Polymarket between 10¢ and 97¢ is accurate to within a few percent of its own price across 33,000 contracts. That's the real headline for anyone using these markets as information rather than as a bet.
- 02Below 10¢, the price is not a probability. It's a probability plus a premium you pay for holding a lottery ticket. On Kalshi that premium is roughly half the contract's value.
- 03The spread is not a rounding error down there. A 3¢ median spread on a 5¢ contract is a 60% round-trip cost before anything else happens.
- 04The mirror trade is not free money. Selling overpriced longshots means posting 98¢ of collateral to win 2¢, for months, with the full loss if you're wrong. That's why the mispricing persists: the trade that would correct it is terrible on capital.
#What I can't tell you
Being straight about the limits, because they're real:
- ▸I can't fully separate venue from subject matter. The within-event test handles it inside each event, but Kalshi and Polymarket still list different kinds of questions over different periods.
- ▸Kalshi's sample is a random slice, not a census. Collecting its full universe would take roughly 940,000 API requests.
- ▸I can't test the capital-cost explanation properly. The idea that longshot prices reflect the cost of tying up collateral is plausible, but my sample is dominated by short-dated contracts where the effect would be a fraction of a percent, too small to detect. That question needs long-dated markets I don't have enough of.
- ▸One result remains unexplained. Polymarket contracts priced 7 days out show a significant miss concentrated in the 40–60¢ range in sports, the opposite end of the price scale from everything else here. I don't know why, and I'm not citing it for or against anything.
Method. 104,372 Polymarket observations across 46,008 markets (Jan 2023 – Jul 2026) and 32,674 Kalshi observations across 21,536 markets (May 2024 – Jul 2026), collected from public APIs. Prices sampled at fixed horizons before resolution, verified at or before cutoff. Excluded: unresolved and voided markets, never-traded default prices, duplicates. All standard errors clustered by event; interval estimates are Clopper–Pearson. Within-event estimates follow Ali (1977) by conditional maximum likelihood, with a cluster bootstrap over events. Analysis run against frozen, hash-stamped snapshots after an earlier round was found to have been computed against files still being written.