Masterclass · Module 08
Module 08

How to tell a real edge from a curve-fit one

Holdout testing, sample size and the multiple-comparisons trap — including the filter that improved 7 configurations out of 7 and still failed.

This is the most useful module here, and it is not about gold. It is about how to tell whether any strategy — yours, ours, or one you paid for — has a real edge or has simply been fitted to the past. Almost everything sold to retail traders fails this.

The problem in one sentence

If you test enough variations, some will look excellent purely by chance, and you will pick those — because picking the best-looking one is exactly what testing feels like it is for.

Test forty variations of a coin-flip strategy and about two will show a profit factor above 1.05 on noise alone. Nothing in the result itself tells you which kind you are holding.

The fix: a holdout

Split your data before you start. Use the first portion to make every decision — settings, filters, thresholds, all of it. Then run the finished thing once on the portion you never looked at.

If the edge is real, it shows up in both. If it was fitted, it evaporates in the second. The discipline is not the split — it is that you only get one look, and you accept the answer.

Our standard
A strategy must clear the bar twice: on the full sample, and on a 30% holdout it was never tuned against. In-sample performance alone rewards whoever searched hardest.

A worked example of it saving us

We tested adding a volatility filter — skip trades when the market is unusually quiet. On the training data it improved profit factor in seven configurations out of seven. Seven for seven. Every variation we tried got better.

That consistency is exactly what a real effect looks like, and we were ready to ship it.

Then we ran it on the holdout. It made things worse: profit factor fell from 1.35 to 1.29 on the data it had never seen.

Seven for seven was a mirage. We dropped the filter.

Sit with that for a second. If we had tested it the way most strategies are tested — try it, see improvement, adopt it — that filter would be in the live system right now, quietly making it worse, and we would be citing "improved in 7 of 7 tests" as evidence in our marketing. It would have sounded rigorous. It was wrong.

The three questions to ask any strategy

  1. How many trades? Under 30 proves nothing in either direction. Under 100 is a small sample and should be described as one. Ours is 123, and we say so every time.
  2. Was it tested on data used to build it? If there is no holdout, the number is a description of the past, not a prediction.
  3. How many variations were tried before this one? "We found a setting that works" means something very different after two attempts than after two hundred.

If a course, signal service or EA cannot answer all three, that is your answer.

Costs are part of the test

One more way results lie. We tested our strategy on USDJPY and it looked tradeable — profit factor 1.12. Then we corrected the cost model to include commission properly. FX moves roughly $0.61 per point per lot, so a $7 round-trip commission is about 11 points, not the 2 that a spread-only model implied.

With the correct cost, that same strategy scored 0.92. A loser. The difference between a product and a mistake was one arithmetic error in the cost model.

What to take from this module

Educational content · not financial advice · trading gold carries substantial risk of loss · past and hypothetical performance never guarantees future results. Risk disclaimer