What a grade has to earn
0xinsider assigns every tracked Polymarket wallet a grade from S down to F, computed from that wallet's own settled history across six metrics with Bayesian shrinkage, so one lucky market does not produce an S. A grade like that is easy to compute and easy to fool yourself with. It only means something if it says something about the trades a wallet has not made yet.
So here is the test. Take every buy of $10,000 or more on Polymarket between June 1 and September 11, 2026, on a market that settled after the trade cleared. For each one, look up the grade that wallet already held on the day it traded, before anyone knew how the market would resolve. Then check the settled outcome. That is 67,531 trades, 4,313 wallets, 11,370 settled markets and $2.10 billion in notional.
The grade separates. Wallets graded S, A or B beat the price they paid. Wallets graded D or F lost to it. Wallets with a C, and wallets with no record yet, came out flat.
Why win rate answers nothing here
A wallet that only buys contracts at 85 cents wins about 85% of the time and has demonstrated no skill whatsoever. Win rate measures which prices someone likes. Report it alone and every favorite-buyer on the platform looks like a genius.
The measure that survives is the gap between how often the chosen side won and the price paid for it. Buy at 65 cents and win 66.5% of the time and you were 1.5 points better than the market that sold to you. Buy at 58.7 cents and win 56.6% of the time and you were 2.1 points worse. That gap is what a grade has to predict, and it is the number this study reports.
It is the same logic behind calibration edge on every trader profile. Paying less than the eventual frequency is the skill. Backing the obvious side is not.
The result
Wallets graded S, A or B took 24,610 of the large buys in the window, across 1,207 wallets and 5,946 settled markets, worth $717.3 million. They won 66.5% of the time at an average price of 65.0 cents. That is an edge of 1.57 points, and 2.22% on a dollar-weighted basis.
Wallets graded D or F took 17,106 buys across 1,557 wallets and 4,476 markets, worth $520.2 million. They won 56.6% of the time at an average price of 58.7 cents, an edge of minus 2.06 points. Wallets graded C landed at plus 0.53 points and wallets with no grade yet at minus 0.13, both close enough to zero to be indistinguishable from the market.
Confidence intervals are bootstrapped over markets rather than over trades, because four wallets buying the same side of the same market are one observation and not four. Treating them as independent is the standard way to manufacture a result that is not there. On 5,000 resamples the S, A and B interval runs from plus 0.32 to plus 2.82 points and the D and F interval from minus 3.61 to minus 0.45. Both clear zero. The C interval and the ungraded interval both straddle it.
Broken out by letter: S posted plus 1.32 points across 11,230 trades from 133 wallets, A plus 1.53 across 4,984 trades from 400 wallets, B plus 1.94 across 8,396 trades from 941 wallets. D posted minus 4.32 across 2,598 trades and F minus 1.65 across 14,508. The order inside those groups does not run the way the letters do, and B beating S is not evidence that B is the better grade. At these sample sizes those differences sit inside the noise. The honest reading is that the grade separates the top three letters from the bottom two and does not finely rank within them.
Whether this is just favorites winning
The obvious objection is that good wallets buy favorites, favorites win, and the grade is measuring nothing beyond a taste for high prices. Splitting the same trades by the price paid settles it.
Under 20 cents, S, A and B wallets posted plus 3.18 points against plus 1.92 for D and F. Between 20 and 40 cents, plus 0.93 against plus 0.01. Between 40 and 60 cents, plus 0.42 against minus 1.60. Between 60 and 80 cents, plus 3.27 against minus 1.93. At 80 cents and above, plus 1.35 against minus 5.27.
The graded cohort beats the ungraded one in all five buckets, which rules out the favorite-longshot explanation. The gap is widest at the top of the price range: a well-graded wallet paying 87 cents is right more often than a poorly graded wallet paying the same 87 cents, by more than six points. Expensive contracts are where a bad read costs the most, and it shows.
Where it works and where it does not
Within the S, A and B cohort, and counting only categories with at least 300 trades, politics posted plus 6.94 points across 593 trades, esports plus 2.85 across 2,185, soccer plus 1.77 across 15,527, tennis plus 0.44 across 2,692, and baseball minus 1.10 across 2,557.
Soccer is 63% of the graded sample, so the headline number is mostly a soccer number. Baseball runs the other way, and politics runs far ahead on a sample small enough that one busy week could move it. Read the category splits as a map of where to look next, not as five separate findings.
What this study does not show
The edge is measured at each trade's own price against the settled outcome. It ignores fees and slippage, and it treats every position as held to resolution. A wallet that bought at 60 cents and sold at 75 before the market resolved is scored on the settlement it never waited for, which is not what it made.
Grades are computed from a wallet's own past resolved trades. The lookup is point-in-time, so no trade is ever scored by a grade that saw its own market resolve, but this is not a held-out universe in the strict sense.
The scored population thinned after July, from about 18,500 wallets a day in June and July to about 3,900 in August and September, so the later slices carry more noise than the earlier ones. Only settled markets are in the sample, which tilts it toward shorter-dated events, and sport is most of that.
And 1.57 points is a small edge. That is the size you should expect from a market that mostly works, and it is the reason to trust the number rather than a reason to dismiss it. A study of this kind that reported ten points would be describing a bug in its own method. None of this says an S-graded trade is worth copying. It says the grade carries information about who tends to be right, which is a different and smaller claim.
How to check it yourself
The sample is every Polymarket buy of $10,000 or more between June 1 and September 11, 2026, at a price between 2 and 98 cents, on a market whose resolution timestamp falls after the trade. Buys only, because a buy at a price on a named outcome is a clean directional bet and a sell is not.
Grade assignment takes the most recent daily ranking dated on or before the day of the trade. The window starts on June 1 because grade coverage widened that month, from about 238 wallets scored per day in May to 18,464 in June. Any version of this study that reaches back further reads an absent grade as a fact about the wallet when it is a fact about the pipeline, and an earlier pass of this analysis did exactly that and produced a March figure of minus 11 points that meant nothing.
Every number on this page comes from one run of the same set of queries against production, at 02:25:45 UTC on September 12, 2026. The grading method behind the letters is documented on the transparency page, and the per-wallet inputs are visible on every trader profile.
Expect your own counts to come out higher than ours. Every trade in the sample sits on a market that has already settled, so the universe grows each time an open market resolves. Three runs of the unchanged query on the day this was published, about ninety minutes apart, returned 67,516, then 67,517, then 67,531 trades. Bounding on the resolution timestamp does not freeze it either, because rows gain an outcome after the fact carrying a timestamp that predates the bound.
The answer held still while the counts moved, which is the part worth knowing. The S, A and B edge came out at 1.57 points in all three runs. D and F came out at minus 2.05, minus 2.05 and minus 2.06. The graded cohort beat the poorly graded one in all five price buckets every time. A reproduction that lands within a few hundredths of a point on a slightly larger sample is the study reproducing, not disagreeing.
Common questions
Do Polymarket wallet grades actually predict outcomes?
Across 67,531 Polymarket buys of $10,000 or more between June 1 and September 11, 2026, wallets graded S, A or B won 66.5% of the time at an average price of 65.0 cents, beating the price they paid by 1.57 points with a market-clustered 95% confidence interval of plus 0.32 to plus 2.82. Wallets graded D or F lost to the price by 2.06 points. C wallets and wallets with no record yet came out flat.
Is the study affected by lookahead bias?
No. For each trade the analysis uses the most recent grade dated on or before the day that trade cleared, so a grade computed after a market resolved is never used to score a trade in that market. Grades are still derived from a wallet's own earlier resolved history, so this is point-in-time rather than a fully held-out universe.
Are the graded wallets just buying favorites?
No. Splitting the same trades into five buckets by the price paid, S, A and B wallets beat D and F wallets in every bucket. The gap is widest above 80 cents, where the graded cohort posted plus 1.35 points and the poorly graded cohort minus 5.27.
Is a 1.57 point edge large enough to trade on?
It is small, and that is what an efficient market should produce. It says the grade carries information about which wallets tend to be right. It does not say that copying an S-graded trade is profitable after fees, slippage and the timing of your own entry, and this study does not test that.
Why not just use win rate to rank traders?
A wallet that only buys contracts at 85 cents wins about 85% of the time without demonstrating any skill. Win rate measures which prices a trader likes. The measure that separates skill from taste is the gap between how often the chosen side won and the price paid for it.
