Every prediction market ships a leaderboard ranked by profit. None of them check whether the leaders are actually good or just ran hot, so the whole ecosystem treats "up money" as "skilled." Over a few dozen bets those are not the same claim.
I have been building Convexly, an independent skill-vs-luck audit for Polymarket wallets, to put a number on that gap. Here is the part that keeps surprising people.
Take the top 50 wallets by profit. Read only their resolved positions. Compute realized edge with a 95% confidence interval from a clustered bootstrap, so 50 bets on one election night do not count as 50 independent data points. Then hold the whole set to a false-discovery-rate bar, which keeps a handful of lucky-looking records from sneaking through when you test many wallets at once. Almost none of the profit leaders clear it. Being up money is mostly a small-sample story, and the wallets that do clear the bar are usually not the ones on top.
A few build decisions that fell out of this:
I would rather be corrected than confident. If you follow prediction markets and think the approach is wrong somewhere, I want to hear it. The analyzer is free and takes a wallet address with no signup, and we are on Product Hunt today if you want to poke at it.
Happy to get into the stats in the comments.
What I found most interesting is that you're measuring the confidence of the conclusion, not just the outcome.
A lot of products stop at ranking performance. You're asking whether the available evidence has actually earned the right to call someone skilled. That's a much higher standard, and a very different product.
Exactly. Raw PnL tells you what happened, not whether it was earned. The FDR bar is the part doing the work: test 50 "winners" at once and some clear any skill threshold by luck alone, so you have to correct for how many shots you took. After that, most of the top 50 don't survive. "Up money" and "skilled" are just different claims, and almost nobody separates them. Appreciate you reading it that closely.
Interesting.
Reading your reply made me think less about the statistical correction itself and more about what changes once a product starts distinguishing between evidence and conclusions instead of treating them as the same thing.
I don't think I can explain that line of reasoning properly in a thread without oversimplifying it.
If you're interested, what's the best email to reach you on?
research@convexly.app comes straight to me. Curious where your thinking goes.
Thanks! I’ve just sent it over.
Looking forward to hearing your thoughts whenever you have a chance.