Measurement · 25 August 2026
Polymarket's biggest wallets aren't traders.
They're market makers.
We built a tracker on the premise that some whales are worth following. Three data sources later — two of them biased in opposite directions, neither saying so — the answer turned out to be worse than "it doesn't work". They have no edge on the trade, and all 25 of the largest collect $4.2M in exchange rebates for posting liquidity. A copier gets the bet and none of the income.
We spent three months building a tool that watches large Polymarket wallets, analyses each position with an LLM, and pushes a BUY / WATCH / SKIP verdict to Telegram. The premise was the same one every tracker in this category runs on: some whales are better than others, and following the good ones beats following the bad ones.
We finally measured the premise. It is false — not unproven, but unmeasurable in principle at the volumes anyone in this category has. Getting there was harder than the result suggests, and the hard part was the data.
Three sources, two of them lying
The obvious way to study whales is to read Polymarket's public trade feed and keep the rows belonging to wallets you care about. That is what we did for three months. It is wrong in a way the feed never mentions.
| Source | Edge vs price | What it gets wrong |
|---|---|---|
Global /trades feed | +0.51 pp | Taker side only |
/closed-positions | +5.65 pp | 14% of losses missing |
| /activity | +0.41 pp | — |
All three measured on the same wallets on 21 August, so the rows compare sources rather than dates. Findings below use the current snapshot.
The global feed reports a trade from the taker's side only. When
the wallet you are watching is the maker — a resting order someone else
crossed into — the trade is attributed to the counterparty and never reaches you.
For one of our wallets, takerOnly=true returned 500 fills spanning
14.3 hours against 500 in 1.7 hours without the filter: only about
11% of its activity was as taker. On top of that, the feed serves
a cache that advances in five-minute jumps, so a page narrower than 300 seconds
physically cannot see everything — we were capturing roughly
30% of even the taker fills.
The closed-positions endpoint drops losers. It looked like the authoritative fix — Polymarket's own record of every position, with its own P&L — and it said whales were up 10.8% with a +5.65 point edge over price. That is far too good, so we tested it: for 300 markets where we already knew the outcome from the exchange's own resolution flag,
winning positions 237 / 237 present 100% losing positions 208 / 242 present 86%
A winning position gets redeemed, so it closes. A losing position is worth nothing and 14% of them simply never become "closed". For that wallet the true win rate is 50.8%; dropping one loss in seven lifts it to about 54.5%. The entire "edge" was the missing losers.
The activity endpoint has neither problem. It is per-wallet, so makers are included (549 rows/hour against 48 with the taker filter), and outcomes come from the exchange's resolution flag rather than from what got redeemed. Everything below uses it: 4,833 resolved positions, 70 wallets, $39.1M staked.
Finding 1: whales have no edge over the price they pay
Under the null hypothesis that the market price is the probability, a position bought at 0.60 should win 60% of the time.
4,833 positions · average entry 0.5193 · actually won 51.36% edge: -0.58 percentage points cluster-robust 95% CI: -1.67 to +0.52 (not significant)
Zero, within the error bars. Whales pay fair prices. This is the same answer our separate study of market calibration gave, from a completely different direction.
Dollar-for-dollar they still lose: -5.33% before fees, -7.54% after a taker's fee. The typical position is a coin flip at a fair price; the largest ones lose. That gap between counting positions and counting dollars shows up in every source we tried.
Finding 2: they are not trying to win the bet
A trader with no edge is a puzzle. A market maker with no edge is a business. We looked at what these wallets are actually paid, and the puzzle dissolved.
Polymarket pays three things that have nothing to do with being right. Maker rebates return 15–25% of collected taker fees to whoever posted the liquidity that got taken. Taker rebates, live since 28 May 2026, run on volume tiers from 3% at $2,000 of 30-day weighted volume to 50% at $10M. And liquidity rewards pay for quoting at all.
All 25 of the largest wallets on our list collect them. Every one.
| Payment type | Across 25 wallets |
|---|---|
| Taker rebates | $1,879,355 |
| Maker rebates | $1,791,930 |
| Liquidity rewards | $547,017 |
| Total | $4,218,302 |
One wallet takes $5,035 a day in taker rebates and $5,042 a day in maker rebates, against total account equity of $1.36M. Another has collected $1.45M. The income is the flow, not the direction — which is exactly why the edge on the trade measures zero. They are not guessing, and they do not need to.
We checked what the rebate is a percentage of, because the docs give the tier formula and not the base. Summing the taker fees five wallets paid over the same window as their payouts:
fees $125,348 -> rebate $34,824 ratio 0.28 (Platinum, 32%) fees $272,086 -> rebate $73,534 ratio 0.27 (Platinum) fees $449,373 -> rebate $174,810 ratio 0.39 (Diamond, 44%) fees $331,666 -> rebate $142,908 ratio 0.43 (Diamond) fees $208,195 -> rebate $58,115 ratio 0.28 (Platinum)
It is a share of fees paid, at the tier rate. So a whale's effective taker cost is 56–73% of the posted rate. A reader copying an alert does $2,000 a month if they are keen, sits at tier zero, and pays 100% — while posting no liquidity and earning no rewards.
That is the whole answer. Copying a whale copies the half of their book that does not make money, at a worse fee, without the half that does. It is not that following them fails to work. It is that the thing they are paid for is not the trade.
Finding 3: you cannot rank these wallets, and the reason is arithmetic
The obvious next move is to find which wallets are good. We ran a null test per wallet — 3,000 simulations each, every position winning with probability equal to its entry price — on the 54 addresses with at least $20k staked and at least 10 resolved positions.
beat the market at p < 0.05: 1 of 54 (expected by chance: 2.7)
One, against the 2.7 you would get from noise alone — fewer than chance, not more. Nothing here survives a correction for having tested 54 wallets, and nothing needs to.
That is not about the sample being small. To detect a 5-percentage-point edge at 80% power you need roughly 620 equal-weight positions. Whales do not bet equal weights — they bet $36 and then $740,000 — and once you account for that, the usable sample collapses:
| Effective sample (Kish) | Positions |
|---|---|
| Best-sampled wallet | 115 |
| Median wallet | 12 |
| Wallets reaching 400+ | 0 of 54 |
The median wallet is worth 12 independent bets. Every product that ranks traders by track record — ours included — is ranking noise at these volumes. That is not a criticism of anyone's engineering. It is a property of unequal bet sizing that no amount of data collection fixes quickly, because the same wallets keep making the same lopsided bets.
Finding 4: most of them lose the bet, and none of them provably deserve to
Of the 61 wallets with at least $20k staked, 36 are in the red and the median return is -7.6%.
An earlier version of this page added that the five largest wallets by volume "hold a third of the money and are down 8.1%." Four days later the same measurement reads +6.8%, and only three of those five wallets are still in the top five. Two esports positions resolved in between. We are leaving the correction visible because it is Finding 3 happening in real time: if the aggregate of the five biggest wallets on the venue swings fifteen points in four days, no ranking built on numbers like these means anything.
But "most whales lose" and "this whale is bad" are different claims, and only the first one holds. Not one of the 54 is distinguishable from bad luck in either direction. A wallet down 40% on sixteen bets at an average price of 0.40 is an ordinary run, not a verdict on the person.
What we got wrong
Five results came out wrong in two days. Three were caught by asking is this number even possible?, one by reading documentation instead of our memory of it, and one only by auditing code that had already produced a published number.
1. The fee, applied to the wrong denominator
Polymarket charges takers shares × rate × price × (1 − price). We had
it as rate × (1 − price) per share — the right expression
for a fraction of stake, applied to the wrong quantity. That overstates
the fee by 1/price. Caught while rewriting a tweet, by reading the
fee docs.
2. Shares derived from an averaged price
A whale fills one leg many times at different prices. We summed the dollars, took the mean price, and divided to get shares. A fill at 0.06 buys a pile of shares while barely moving a mean anchored at 0.60. It understated shares by 5.2% and the gross return by 2.3 points of stake.
3. An anchor that depended on the outcome
Measuring price calibration, we sampled each market's price N hours before
min(endDate, closedTime). A market resolving YES closes early, the
moment the event happens; one resolving NO runs to its deadline. The measurement
time was a consequence of the result, and the answer came out exactly
backwards. Fixed by anchoring on a calendar date instead.
4. Comparing dust to real bets
Size quintiles showed a dramatic effect: in the 0.30–0.40 band the smallest 20% of positions won 37.6% and the largest 20% won 22.0%. The smallest quintile had a median size of $2–7. On real sizes the effect vanishes.
5. Paginating an endpoint that sorts by profit
/closed-positions defaults to sortBy=REALIZEDPNL
descending. We pulled 3,000 rows for one wallet and computed a
+84% return on $90.4M. Those were the 3,000 most profitable
positions out of more than a hundred thousand. Sorting the other way, the same
50 rows run to -$208,060. This one happened during the audit that was
supposed to catch things like it.
What survived
One thing did. Our own verdicts discriminate, slightly. Stratified by price band, so price cannot explain it:
BUY vs WATCH: +7.7 points (permutation p = 0.013) BUY vs rest: +5.8 points (permutation p = 0.028)
But BUY's absolute edge is +2.0 points and not significant (p = 0.21). At an average entry of 0.541 you need 55.4% to break even after a taker's fee and we are at 56.1% — which sounds like a result and is not one: proving an edge that small would take roughly 9,950 BUY calls, more than a decade at our rate.
So the honest summary of our own product: it is measurably better at saying don't than at saying do, and the do side is indistinguishable from breaking even.
Caveats
- The verdict result is marginal: two tests near p = 0.02–0.03, and a sign test across price bands does not reach significance. Several tests were run. We would not call it proven.
- Positions are sampled across each wallet's history at spread offsets, not exhaustively. One wallet alone has traded 112,491 markets.
- Effective sample size (Kish) approximates power. It is right to an order of magnitude, not to the second digit.
- 70 wallets over one venue. Whether any of this generalises, we do not know.
Reproducing this
Every number above comes from public endpoints — wallet activity, market resolution flags, order books and price history. No private data is needed beyond our own alert history. Numbers were last measured on 25 August 2026 and move as positions resolve — see Finding 4 for how much.
We build WhaleSense, the tracker described above. It is still running, with the claims corrected.