Holdout testing, sample size and the multiple-comparisons trap — including the filter that improved 7 configurations out of 7 and still failed.
This is the most useful module here, and it is not about gold. It is about how to tell whether any strategy — yours, ours, or one you paid for — has a real edge or has simply been fitted to the past. Almost everything sold to retail traders fails this.
If you test enough variations, some will look excellent purely by chance, and you will pick those — because picking the best-looking one is exactly what testing feels like it is for.
Test forty variations of a coin-flip strategy and about two will show a profit factor above 1.05 on noise alone. Nothing in the result itself tells you which kind you are holding.
Split your data before you start. Use the first portion to make every decision — settings, filters, thresholds, all of it. Then run the finished thing once on the portion you never looked at.
If the edge is real, it shows up in both. If it was fitted, it evaporates in the second. The discipline is not the split — it is that you only get one look, and you accept the answer.
We tested adding a volatility filter — skip trades when the market is unusually quiet. On the training data it improved profit factor in seven configurations out of seven. Seven for seven. Every variation we tried got better.
That consistency is exactly what a real effect looks like, and we were ready to ship it.
Then we ran it on the holdout. It made things worse: profit factor fell from 1.35 to 1.29 on the data it had never seen.
Seven for seven was a mirage. We dropped the filter.
If a course, signal service or EA cannot answer all three, that is your answer.
One more way results lie. We tested our strategy on USDJPY and it looked tradeable — profit factor 1.12. Then we corrected the cost model to include commission properly. FX moves roughly $0.61 per point per lot, so a $7 round-trip commission is about 11 points, not the 2 that a spread-only model implied.
With the correct cost, that same strategy scored 0.92. A loser. The difference between a product and a mistake was one arithmetic error in the cost model.