How to Compare Trading Bot and Strategy Performance Fairly

Compare strategies on the same window using return, drawdown, return÷drawdown, Sharpe and behaviour in sell-offs, with a worked 10-year backtest example.

League Basics · Updated 2026-10-08

Start with the same window and the same rules

Two performance numbers can only be compared if they were measured the same way. That means the same start and end dates, the same starting capital assumptions, the same treatment of fees and the same trading mode. A strategy that started just after a crash will look brilliant next to one that started just before it, even if the rules are no better. A paper result next to a live result is not a comparison at all.

On the strategy pages, every backtest is re-run on roughly ten years of daily data from October 2016 to October 2026, which removes the most common source of unfair comparisons. In Live Paper Trading, bots that share a measurement window can be compared with each other, but a four-week window is far too short to say much about either.

Return and drawdown belong together

Total return and CAGR, the compound annual growth rate, tell you how much a strategy made. Maximum drawdown tells you the worst fall from a previous high along the way. Neither is meaningful without the other. A strategy that earns 25% a year with a 50% drawdown asks you to sit through losing half your money at some point; one that earns 15% with a 15% drawdown never does.

The maths of losses makes drawdown more important than it first looks. A fall of X% needs a gain of 1 ÷ (1 − X) − 1 to get back to even. A 15% drawdown needs about +17.6%; a 50% drawdown needs +100%. Deep drawdowns do not just hurt, they take a long time to repair, and many people abandon a strategy before that repair happens.

Return ÷ drawdown, Calmar and Sharpe

A simple way to put the two together is to divide annual return by the size of the maximum drawdown. The Calmar ratio is usually defined as CAGR divided by maximum drawdown, and it answers a plain question: how much yearly return did I get for each unit of worst-case pain? A ratio above 1 means the strategy earned more per year than its deepest fall.

The Sharpe ratio looks at a different kind of risk. It divides average excess return by the volatility of returns, so it penalises a bumpy ride even if there was never one large crash. Sharpe is useful for comparing day-to-day smoothness; return ÷ drawdown is useful for comparing worst cases. Read them together, and remember that both become unreliable on short samples. A Sharpe annualized from a few weeks of daily data can look enormous and mean very little.

Why win rate misleads

Win rate is the share of trades that closed with a profit. It is easy to understand and easy to misuse. A strategy that wins 90% of the time with an average gain of 1% and loses 10% of the time with an average loss of 10% makes +9% and −10% over ten typical trades, a net loss before fees. Grid and mean-reversion strategies often show high win rates for exactly this reason: they take many small profits and occasionally hold a large loss when price runs away. Always pair win rate with the average win, the average loss and the maximum drawdown.

Worked example: Growth-Value Style Selection vs Kaufman Efficiency Index

Here are two strategies from the backtests re-run on 7 October 2026, both covering 24 October 2016 to 5 October 2026. Growth-Value Style Selection shows a CAGR of 29.54%, a maximum drawdown of −34.97% and a Sharpe of 0.97, for a total return of 1211.9%. Kaufman Efficiency Index shows a CAGR of 25.09%, a maximum drawdown of −48.62% and a Sharpe of 0.86, for a total return of 826.6%. Growth-Value Style Selection is an aggressive-tier strategy: it can hold 3x leveraged ETFs (SPXL, TQQQ, UPRO), it was selected from many backtested candidates so luck cannot be ruled out, and it has no live or forward record yet.

Here the headline numbers point the same way: Growth-Value Style Selection has the higher return and the shallower drawdown. Return ÷ drawdown makes the gap concrete: 29.54 ÷ 34.97 = 0.84 against 25.09 ÷ 48.62 = 0.52. Recovery maths sharpens the point: getting back from −34.97% needs +53.8%, while getting back from −48.62% needs +94.6%. Kaufman Efficiency Index's deepest drawdown ran from a peak on 26 January 2018 to a trough on 14 August 2019, while SPY returned +1.9% over the same stretch, and the strategy regained its high on 19 February 2020. Growth-Value Style Selection's deepest fall ran from 26 January 2018 to 24 April 2018 and was recovered on 8 January 2020: a short fall followed by a long climb back.

Calendar years add another angle. Growth-Value Style Selection beat SPY in 5 of the 9 full years and Kaufman Efficiency Index in 5. Kaufman Efficiency Index's best year was 2017 at +64.1% against SPY's +21.7%. Growth-Value Style Selection's best year was 2020 at +103.6% against +18.3%. A higher ten-year CAGR still left several years behind the index.

The SPY sell-offs show how differently the two behave. SPY fell more than 10% three times in the window. From 20 September 2018 to 24 December 2018, while SPY moved −19.4%, Growth-Value Style Selection returned −19.5% and Kaufman Efficiency Index −12.0%. From 19 February 2020 to 23 March 2020 (SPY −33.7%), they returned −14.3% and −21.7%. From 3 January 2022 to 12 October 2022 (SPY −24.5%), they returned −22.8% and −16.3%. Neither strategy was consistently the better brake: which one held up better changed from one sell-off to the next.

Better headline numbers do not settle the comparison on their own. Check how each result was produced (leverage, how many candidates were tried, whether there is any live record) and how each behaved in the sell-offs you care about. Which fits depends on how large a fall you could actually sit through.

Regimes and sample length

Strategies are built for conditions. Trend-following and momentum rules tend to do well in persistent trends and poorly in choppy, range-bound markets; grid and mean-reversion rules tend to do the opposite. A comparison over a period that contains only one kind of market mostly tells you which strategy suited that market. That is why ten years of data, with several sell-offs and recoveries, says far more than a few months, and why a few weeks in Live Paper Trading says very little on its own. Even ten years is one path through history; the next decade will not repeat it.

A short comparison checklist

Before comparing, confirm that the window, the mode and the cost assumptions are the same. Then read return and maximum drawdown together, compute return ÷ drawdown, and check the Sharpe for smoothness. Look at calendar years to see whether the result came from one exceptional year, and look at behaviour during broad sell-offs to see whether the rules actually protected capital when it mattered. Treat win rate as a supporting detail, never the headline. Finally, ask how long the record is and how many different market conditions it has been through.

Risk note

Backtests and paper records describe the past under stated assumptions. They are not forecasts, they leave out some real-world costs, and the strategy with the best past numbers is not guaranteed to do best next. Nothing here is investment advice. Any real trading can lose money, and a strategy with a deep historical drawdown can produce an even deeper one in the future.

See it in the data

Related guides

Metrics explained with real backtests

Educational content, not investment advice. Guide content and league rankings are independent of any exchange or sponsor relationship.