FINAXIONResearch
Finaxion Research

How we test a strategy (and why you can't accuse us of cherry-picking)

This document was published before the first verdict. It is the contract. Every strategy in the lab goes through exactly this process, with the same rules, the same data and the same costs. No exceptions for sympathy or fame.

In one sentence

We take the rule as it is sold, turn it into something mechanical a machine can execute without interpreting, run it over each asset's full available history (up to six decades) that the strategy never saw, charge it real commissions, and check whether the result survives when we nudge the numbers or look at years that were not used to tune it.

1. The claim is quoted, not interpreted

Every page opens with the promise, in quotation marks, with a source. Below it, exactly what we tested: the mechanical version of that promise. If the promise is vague ("buy when smart money accumulates"), we say which approximation made it executable, and we label it as an approximation. Testing one thing and claiming to have tested another would be this lab's original sin.

When a school is discretionary (ICT, Wyckoff, Elliott), what gets tested is the most-cited, mechanisable version, not any specific trader. Nobody can say "I do it differently": true, and that is why the exact rule is on display.

2. The data

3. The costs

Every trade pays commission and slippage per side, with a different cost per asset class, because a single number would be unfair in both directions:

Class Per side Round trip Reference
Stocks and index 10 bps 20 bps retail investor at a modern broker, liquid stocks
Forex 1 bp 2 bps majors at a retail broker (spread ≈ 0.5-1 pip)
Crypto 15 bps 30 bps exchange taker fee + slippage

Strategies that trade a lot pay a lot: that is real life, and it is why so many rules "work" on paper and not in the account. Intraday, with thousands of trades, cost is almost always what decides.

No leverage. No shorting unless the rule demands it. No stop-loss added by us unless the rule includes one: a stop changes the strategy.

4. The yardstick: buy and hold

The question is not "did it make money?". Over bullish decades almost everything makes money. So every market is measured with two yardsticks: buying and holding that same asset (did the rule beat doing nothing?) and buying and holding the S&P 500 over the same window (was it worth more than the index?). The first decides the verdict; the second is shown in every table, in CAGR points per year, from 1993. The question is did it make more than doing nothing? Every strategy is compared with buying the index (S&P 500, via SPY) on day one and never touching it. The key figure on every page is "Vs. index": the difference in total-return points between the strategy and that boring alternative.

5. Six arenas

A timing rule tested only on the S&P 500 in a bull decade faces the hardest test there is, and that is not where people use it. They use it on NVDA, on bitcoin, on EUR/USD. So every strategy runs in five arenas, and on each asset it is compared with buying and holding that same asset:

  1. Indices and ETFs: S&P 500, Nasdaq-100, Russell 2000, emerging markets, gold and long bonds, one by one. The S&P 500 (SPY) is also the reference for the headline chart, the robustness suite and the parameter neighbours.
  2. Balanced portfolios, the way the platform backtests: the Dow 30 (at most 10 positions) and the Nasdaq-100 (at most 20), equal weight. They test whether the rule helps to pick as well as to time.
  3. Stocks, one by one: every member of the Dow 30 and the Nasdaq-100, each against holding that same stock.
  4. Forex: the eight majors against the dollar (EUR, GBP, JPY, AUD, CAD, CHF, NZD and MXN). Indicators that need volume are marked not applicable here (the FX market reports no consolidated volume).
  5. Crypto: bitcoin, ether and the main alternatives, each from its first bar (seven days a week).

Over 130 markets per strategy. The page shows the result on every one.

An honest word on survivorship. The index members we use are today's: the companies that survived and thrived. That makes buy and hold a higher bar than it was at the time — the bias works against the strategies, not for them. A rule that beats the index despite that handicap is more credible, not less. We say it so nobody has to discover it.

5-bis. And intraday?

People do not trade these rules on daily candles: they trade them on 15-minute and 1-hour bars. So every strategy is also tested on 5-minute, 15-minute, 1-hour and 4-hour bars, over all the intraday history that exists for each asset: stocks and indices since 2016 (the full Dow 30 and Nasdaq-100 plus six ETFs), the eight forex pairs since 2003 and crypto since 2015. (The usual desktop sources reach one or two years intraday, which is not enough for a verdict; the decades come from consolidated market histories.)

Every combination (asset × timeframe) is compared with buying and holding that asset over the same window, with the same per-trade costs. No walk-forward: five anchored folds over a million 5-minute bars would cost hours per strategy to answer something the daily panel already answers. The intraday panel answers one question, and answers it with up to twenty-three years of data: does the rule beat doing nothing on the bars people actually trade? It is shown on every page as an asset × timeframe matrix; it does not change the verdict, which remains daily.

6. The four honesty checks

A pretty curve is not enough. Every strategy goes through:

a) Walk-forward (out of sample)

We split each asset's history into five segments. In each, the rule is evaluated over a period that was not used for anything before. If the strategy wins overall but loses in most segments, what it had was a couple of good years, not an edge. The walk-forward verdict is ROBUST, INCONCLUSIVE or OVERFIT. One honest caveat: a slow rule (the 50/200 cross trades about once a year) leaves a single trade per segment, and nothing can be judged from that. With fewer than three trades per segment on average, the walk-forward is marked INCONCLUSIVE and stops deciding the verdict; the other three checks carry the weight.

b) Deflated Sharpe

The Sharpe ratio measures return per unit of risk. But if you try five variants and keep the best, that best one has a Sharpe inflated by pure selection. The Deflated Sharpe (Bailey and López de Prado, 2014) discounts that effect. We require it to clear 0.95: a 95% probability that the edge is not the prize for having tried several things.

c) Parameter neighbours

If the golden cross works at 50/200 but dies at 40/200 and 60/200, you don't have a strategy: you have a lucky number. We test three to five nearby combinations and count how many also beat the index. Fewer than half = fragile.

d) By halves

We compare the first half of the period with the second. A rule that beat the index from 2016 to 2021 and lost to it from 2021 to 2026 does not "not work": it used to work and stopped. That is a different verdict, and an important one, because many famous rules became popular right when they stopped being useful.

We also run the engine's robustness suite: structural flaw detection on the trade series, trade shuffling (does the result depend on the order trades happened to land in?) and block bootstrap resampling.

7. The result: where it works, not a pass or a fail

A rule that fails to beat buy-and-hold in most markets is not a fraud. Almost no mechanical rule beats everything; what matters is where it does: which asset class, which timeframe, and whether it is traded on a single asset or by buying whichever of a basket is in signal. That is what every page answers first, with the exact count of markets beaten always in view.

The name we give each rule comes from a deterministic function of the same inputs as always. It names the ground; it judges no one:

Name Condition
Works almost everywhere Beats buy-and-hold in at least 60% of the markets tested, and walk-forward is not OVERFIT in at least half of the markets it beats, and the parameter neighbours (on the index) do not flag it as fragile.
Works on its own ground Beats in between 25% and 60% of the markets. The page says which: that is where its real value lives.
Used to work, stopped Beats in fewer than 25% over the full period, but in most of those markets it beat in the first half and not in the second.
Works in few markets Beats in fewer than 25% of the markets tested, without the pattern above. Its ground exists and is published; it is narrow.

Two runs of the same catalogue over the same data produce the same file. Every page carries a receipt: computation date, period, costs and an artifact identifier.

8. What this lab is NOT

9. What you can do with this

Every strategy in the lab also exists inside Finaxion, where you can change the parameters, the universe, the period and the costs, and see what happens to your version. If you think we tested a rule wrongly, write to us with the exact rule you want to see: we run it under the same protocol and publish the result, whatever it is.