Backtest checklist: 12 questions

Data, look-ahead, overfitting, costs and fills: what to check in any backtest, ours included.

Data

  1. 1. Does the period cover several market regimes?

    A few bull-market years say nothing about behaviour in a crash. Look for falls like 2008, 2020 and 2022.

    On this site: Famous method experiments use 26–33 years, mostly from late 1993 (1998 or 2000 for some assets); strategy backtests use about 10 years (2,500 trading days). Year-by-year returns are shown next to SPY.

  2. 2. Was it tested on more than today's survivors?

    Leaving out what disappeared erases those losses (survivorship bias).

    On this site: Strategies ranked from a list of today's survivors were removed in the methodology review. The selection bias of US ETF strategies using today's ETF list is stated in the methodology limits.

  3. 3. Are prices adjusted for dividends and splits?

    Without dividends, stock and bond returns look too low; without splits, fake crashes appear.

    On this site: Published backtests and the famous method experiments use prices adjusted for splits and distributions.

Look-ahead

  1. 4. Is each fill after the decision?

    Deciding on today's close and buying at that same close means knowing the day's move in advance.

    On this site: Strategy backtests decide on closed daily bars and fill at the next day's open; the famous method experiments fill at the next day's close.

  2. 5. Does any signal use same-day information?

    One indicator using the same day's value is enough to inflate a result badly.

    On this site: The October 2026 methodology review removed strategies whose signals used same-day information.

  3. 6. How is the warm-up period handled?

    A 200-day average means nothing until 200 days exist. Filling that gap arbitrarily changes the result.

    On this site: Earlier history is loaded so indicators can be computed, and the strategy holds cash until they can.

Overfitting

  1. 7. Were parameters chosen on the same data the result is shown on?

    If so, the result is not out-of-sample; part of the good number is luck.

    On this site: Some strategy-backtest parameters were chosen on the published period, and the methodology page says so: they are not out-of-sample.

  2. 8. Does it hold up out of sample?

    Good early and broken late usually means fitted to the past rather than a rule.

    On this site: The famous method experiments show CAGR and Sharpe for the last 30% separately, next to SPY.

  3. 9. Is there a forward record after the rules were fixed?

    Only a record made on new data after the rules were frozen avoids fitting the past.

    On this site: Public league strategies build a forward virtual-trading record and are not ranked before 28 days and one fill.

Costs and fills

  1. 10. Are trading costs included?

    A cost-free backtest flatters frequent trading strategies badly.

    On this site: Strategy backtests charge 0.10–0.25% per side; the famous method experiments charge 0.1% per side.

  2. 11. Does it survive double the cost?

    Real costs are often higher than assumed; cost-sensitive strategies flip on small differences.

    On this site: The famous method experiments re-run with double costs and show both.

  3. 12. What is the result compared with?

    A return on its own says little. Compare it with simply holding the market over the same dates.

    On this site: Every backtest is compared with holding SPY over the same dates, with return, drawdown and Sharpe side by side.

Famous method experiments · Methodology

Educational material, not investment advice or a trading signal. Past results do not guarantee future returns.