Far more than most traders use. Thirty trades is within the range of pure luck; a few hundred is the honest minimum for a thin edge, and low win-rate strategies need more still. Our own 173-trade sample is enough to reject a broken idea but not enough to be confident in a profit factor of 1.12.
A strategy with no edge at all will still produce winning runs. Flip a fair coin thirty times and you will regularly see stretches that look like skill. Trading results behave the same way, except the outcomes are unevenly sized, which makes the illusion stronger.
The lower your win rate, the worse this gets. A strategy that wins 25.4% of the time concentrates its profit into a small number of large winners. Whether two or three of those land inside your sample can flip the entire result — which means a thirty-trade test is measuring luck, not the strategy.
We tested on 60 000 bars covering 2.54 years, which produced 173 trades for the sweep strategy. We report its profit factor as 1.12 — and we also say plainly that 173 trades is a modest sample.
Here is the honest reading. It is easily enough to reject the strategies that failed: the opening range breakout lost money across 1005 trades and the momentum strategy across 1434. Large samples losing consistently is strong evidence of no edge. But 173 trades showing a thin positive result is much weaker evidence of an edge.
Rejecting is easier than confirming. That asymmetry is why we describe our surviving strategy as a candidate for forward testing rather than a proven system.
The lower the win rate, the more trades you need before the average means anything:
| Win rate | Rough minimum sample | Why |
|---|---|---|
| 50-60% | ~200 trades | Outcomes are frequent and similarly sized |
| 40% | ~300 trades | Profit concentrates into fewer winners |
| 25-30% | ~400+ trades | A handful of large winners drive the entire result |
| <20% | 500+ trades | Extremely dependent on rare outliers |
A thousand trades from a single quiet month is not a robust sample. You want trades spread across different market conditions: trending and ranging, high and low volatility, different years. Our 2.54-year window contains a strong gold trend, and we cannot claim the behaviour persists in a different regime.
You also need costs modelled. A large sample of cost-free results simply confirms a fantasy with more precision. Every figure we publish has 30 points round-trip spread subtracted from every trade subtracted.
More important than raw count is whether results hold on data the strategy was never tuned on. Walk-forward testing — optimise on one block, test on the next unseen block, roll forward — produces many out-of-sample results instead of one flattering curve.
Our related work ran 32 such folds, of which the sweep passed 21. That is a more meaningful statement than any single backtest number, because each fold was genuinely unseen.
You do not need certainty to act — you need proportionate confidence. With a small positive sample, the correct response is to forward-test at minimum size while you accumulate trades, not to size up because the backtest looked good.
Track your live expectancy against your tested expectancy. If they diverge sharply over a meaningful number of trades, that is signal. If you are eleven losses into a strategy whose worst tested streak was 11, that is noise — see what is a normal losing streak.
Most retail strategy decisions are made on twenty to fifty trades, which is roughly equivalent to deciding a coin is biased after a short run of heads. If you take one thing from this: your sample is almost certainly smaller than you think, and the appropriate response is smaller position sizes rather than more confidence.
No. Thirty trades is well within the range of random variation, especially for low win-rate strategies where a few large winners drive the result.
173 for the strategy we kept, 1005 and 1434 for the two that failed — across 2.54 years of data.
No. It also needs to span different market conditions and to include realistic trading costs.
Out-of-sample validation such as walk-forward testing, where the strategy is measured on data it was never tuned on.