This document was published before the first verdict. It is the contract. Every strategy in the lab goes through exactly this process, with the same rules, the same data and the same costs. No exceptions for sympathy or fame.
We take the rule as it is sold, turn it into something mechanical a machine can execute without interpreting, run it over each asset's full available history (up to six decades) that the strategy never saw, charge it real commissions, and check whether the result survives when we nudge the numbers or look at years that were not used to tune it.
Every page opens with the promise, in quotation marks, with a source. Below it, exactly what we tested: the mechanical version of that promise. If the promise is vague ("buy when smart money accumulates"), we say which approximation made it executable, and we label it as an approximation. Testing one thing and claiming to have tested another would be this lab's original sin.
When a school is discretionary (ICT, Wyckoff, Elliott), what gets tested is the most-cited, mechanisable version, not any specific trader. Nobody can say "I do it differently": true, and that is why the exact rule is on display.
Every trade pays commission and slippage per side, with a different cost per asset class, because a single number would be unfair in both directions:
| Class | Per side | Round trip | Reference |
|---|---|---|---|
| Stocks and index | 10 bps | 20 bps | retail investor at a modern broker, liquid stocks |
| Forex | 1 bp | 2 bps | majors at a retail broker (spread ≈ 0.5-1 pip) |
| Crypto | 15 bps | 30 bps | exchange taker fee + slippage |
Strategies that trade a lot pay a lot: that is real life, and it is why so many rules "work" on paper and not in the account. Intraday, with thousands of trades, cost is almost always what decides.
No leverage. No shorting unless the rule demands it. No stop-loss added by us unless the rule includes one: a stop changes the strategy.
The question is not "did it make money?". Over bullish decades almost everything makes money. So every market is measured with two yardsticks: buying and holding that same asset (did the rule beat doing nothing?) and buying and holding the S&P 500 over the same window (was it worth more than the index?). The first decides the verdict; the second is shown in every table, in CAGR points per year, from 1993. The question is did it make more than doing nothing? Every strategy is compared with buying the index (S&P 500, via SPY) on day one and never touching it. The key figure on every page is "Vs. index": the difference in total-return points between the strategy and that boring alternative.
A timing rule tested only on the S&P 500 in a bull decade faces the hardest test there is, and that is not where people use it. They use it on NVDA, on bitcoin, on EUR/USD. So every strategy runs in five arenas, and on each asset it is compared with buying and holding that same asset:
Over 130 markets per strategy. The page shows the result on every one.
An honest word on survivorship. The index members we use are today's: the companies that survived and thrived. That makes buy and hold a higher bar than it was at the time — the bias works against the strategies, not for them. A rule that beats the index despite that handicap is more credible, not less. We say it so nobody has to discover it.
People do not trade these rules on daily candles: they trade them on 15-minute and 1-hour bars. So every strategy is also tested on 5-minute, 15-minute, 1-hour and 4-hour bars, over all the intraday history that exists for each asset: stocks and indices since 2016 (the full Dow 30 and Nasdaq-100 plus six ETFs), the eight forex pairs since 2003 and crypto since 2015. (The usual desktop sources reach one or two years intraday, which is not enough for a verdict; the decades come from consolidated market histories.)
Every combination (asset × timeframe) is compared with buying and holding that asset over the same window, with the same per-trade costs. No walk-forward: five anchored folds over a million 5-minute bars would cost hours per strategy to answer something the daily panel already answers. The intraday panel answers one question, and answers it with up to twenty-three years of data: does the rule beat doing nothing on the bars people actually trade? It is shown on every page as an asset × timeframe matrix; it does not change the verdict, which remains daily.
A pretty curve is not enough. Every strategy goes through:
We split each asset's history into five segments. In each, the rule is evaluated over a period that was not used for anything before. If the strategy wins overall but loses in most segments, what it had was a couple of good years, not an edge. The walk-forward verdict is ROBUST, INCONCLUSIVE or OVERFIT. One honest caveat: a slow rule (the 50/200 cross trades about once a year) leaves a single trade per segment, and nothing can be judged from that. With fewer than three trades per segment on average, the walk-forward is marked INCONCLUSIVE and stops deciding the verdict; the other three checks carry the weight.
The Sharpe ratio measures return per unit of risk. But if you try five variants and keep the best, that best one has a Sharpe inflated by pure selection. The Deflated Sharpe (Bailey and López de Prado, 2014) discounts that effect. We require it to clear 0.95: a 95% probability that the edge is not the prize for having tried several things.
If the golden cross works at 50/200 but dies at 40/200 and 60/200, you don't have a strategy: you have a lucky number. We test three to five nearby combinations and count how many also beat the index. Fewer than half = fragile.
We compare the first half of the period with the second. A rule that beat the index from 2016 to 2021 and lost to it from 2021 to 2026 does not "not work": it used to work and stopped. That is a different verdict, and an important one, because many famous rules became popular right when they stopped being useful.
We also run the engine's robustness suite: structural flaw detection on the trade series, trade shuffling (does the result depend on the order trades happened to land in?) and block bootstrap resampling.
A rule that fails to beat buy-and-hold in most markets is not a fraud. Almost no mechanical rule beats everything; what matters is where it does: which asset class, which timeframe, and whether it is traded on a single asset or by buying whichever of a basket is in signal. That is what every page answers first, with the exact count of markets beaten always in view.
The name we give each rule comes from a deterministic function of the same inputs as always. It names the ground; it judges no one:
| Name | Condition |
|---|---|
| Works almost everywhere | Beats buy-and-hold in at least 60% of the markets tested, and walk-forward is not OVERFIT in at least half of the markets it beats, and the parameter neighbours (on the index) do not flag it as fragile. |
| Works on its own ground | Beats in between 25% and 60% of the markets. The page says which: that is where its real value lives. |
| Used to work, stopped | Beats in fewer than 25% over the full period, but in most of those markets it beat in the first half and not in the second. |
| Works in few markets | Beats in fewer than 25% of the markets tested, without the pattern above. Its ground exists and is published; it is narrow. |
Two runs of the same catalogue over the same data produce the same file. Every page carries a receipt: computation date, period, costs and an artifact identifier.
Every strategy in the lab also exists inside Finaxion, where you can change the parameters, the universe, the period and the costs, and see what happens to your version. If you think we tested a rule wrongly, write to us with the exact rule you want to see: we run it under the same protocol and publish the result, whatever it is.