Twelve wins out of twenty trades looks like a 60% win rate. It might also be a coin flip having a good week. The gap between those two readings is the single most expensive thing a funded trader can misjudge.
Why a raw win rate lies at small sample sizes
A win rate is just a fraction: wins divided by total trades. The problem is that the fraction carries no memory of how many trades produced it. A trader who is 6-for-10 and a trader who is 600-for-1000 both show “60%,” yet only one of those numbers means anything.
At small sample sizes, randomness dominates. A genuinely break-even system can string together a hot run that looks like a real edge, and a genuinely profitable system can post a losing month that looks like the edge is gone. If you size up after the hot run — which is exactly when most prop traders reach for more contracts — you are scaling a number that hasn’t earned your trust yet.
The raw win rate also hides its own uncertainty. It reports a single point (“60%”) when the honest answer is a range. What you actually want to know is: given what I’ve seen, how low could my true win rate plausibly be?
What a Wilson confidence interval tells you about your edge
The Wilson score interval answers that question directly. Instead of a single win-rate point, it gives you a lower and upper bound on your true underlying win rate at a chosen confidence level (commonly 95%).
Two properties make it the right tool for trade data:
- It shrinks as your sample grows. Twenty trades produce a wide band; two hundred produce a tight one. The interval width is your uncertainty, made visible.
- It behaves sensibly at the extremes. A naive interval can suggest impossible win rates above 100% or below 0% when you’ve had a lopsided run. Wilson pulls the estimate toward the middle in a principled way and stays inside 0–100%, so a 10-for-10 start doesn’t read as “100% forever.”
The number that should drive your decisions is the lower bound. If your point win rate is 58% but the Wilson lower bound sits at 41%, you do not yet have evidence of a positive edge worth scaling. When the lower bound clears the break-even win rate your reward-to-risk requires, that’s a signal built on evidence rather than hope.
This is exactly how Shibiki renders live edge health: each strategy shows its win rate as a Wilson interval that tightens as trades accumulate, so “edge confirmed” is a statistical state, not a gut feeling.
Feeding trades from any platform into one edge model
An edge model is only as good as the trades you feed it, and most traders split their fills across MT5, cTrader, Tradovate or a broker platform depending on the firm. Scoring each platform separately fragments the very sample you’re trying to grow.
The fix is to normalize every fill — entry, exit, size, fees — into one record format and score them together. That’s the design goal behind Shibiki’s auto-journaling: trades land in one edge model regardless of where they executed, so the confidence interval reflects your whole trading, not one account’s slice of it. It also means a strategy you run across several prop accounts is evaluated as a single edge instead of several under-powered ones.
Expectancy and R-multiple as the core inputs
Win rate alone never defines an edge — a 40% system can print money and a 70% system can bleed. The two inputs that actually matter are expectancy and the R-multiple of each trade.
An R-multiple expresses a result in units of initial risk: risk $200, make $400, that’s +2R; hit your stop, that’s −1R. Logging every trade in R makes results comparable across instruments and account sizes. If R-multiples are new to you, start with the R-multiple primer.
Expectancy rolls win rate and average R into one number: the average R you earn per trade over many trades. Positive expectancy is the definition of an edge; the Wilson interval tells you how sure you can be that it’s real. Work through the mechanics in the trading expectancy guide, and put your own numbers in with the expectancy calculator.
Together they give you the full picture: expectancy says how big the edge is, the confidence interval says how confident you should be it exists.
Knowing when your sample is big enough to trust
There is no magic trade count, but there are honest checkpoints:
- Under ~20 trades: treat every conclusion as provisional. The interval is too wide to distinguish skill from variance.
- 20–50 trades: patterns start to firm up. Use the Wilson lower bound, not the point estimate, for any sizing decision.
- 50+ trades of the same setup: the band is usually tight enough to act on — provided the trades were genuinely the same strategy and not a mix relabeled after the fact.
Two rules keep the sample honest. First, don’t pool different setups into one edge score; a scalp and a swing are two samples, not one. Second, let the interval move you before your P&L does — when the lower bound crosses break-even, that’s your cue to add or cut size, ahead of any single day’s result.
Dedicated journals like Edgewonk surface expectancy well; the piece most tools skip is the confidence band around it. Whether you use a spreadsheet, a journal, or Shibiki’s live edge health, the discipline is the same: never scale a number until its lower bound has earned it.
Related: Trading expectancy · R-multiple · Expectancy calculator