What Is an AI Trading Bot Leaderboard? How the League Score Works

How a paper-traded bot leaderboard is built, how the 0–100 league score is calculated with a worked example, and what rankings cannot tell you.

League Basics · Updated 2026-10-08

What a trading-bot leaderboard is

A trading-bot leaderboard is a table that runs several automated strategies under the same rules and sorts them by a common score. The idea is simple: if every bot starts from the same conditions, is priced from the same market data and is measured with the same formulas, the differences you see are more likely to come from the strategies themselves than from the way each result was reported. A single screenshot of a winning bot tells you almost nothing, because you do not know how long it ran, what it risked or how many other bots were quietly switched off. A leaderboard tries to fix that by putting every record side by side, losers included.

That does not make a leaderboard a list of recommendations. It is a measurement tool. Its value depends entirely on how it is built, how the score is calculated and how honestly it shows the bots that are not doing well. The rest of this guide explains each of those pieces for Live Paper Trading, then walks through the score for one real row.

How this league is built

Every public entry in Live Paper Trading is a combination of a strategy template and a market. The templates are familiar families of rules: grid, DCA, momentum, mean reversion, regime switch and pairs spread. The markets are USDT pairs of the kind listed on Binance- and Bybit-style exchanges, such as ETH/USDT, SOL/USDT and BTC/USDT, plus a BTC/ETH spread for the pairs template. So a row labelled DCA Bot on ETH/USDT is not the same thing as a DCA Bot on SOL/USDT, and results should always be read together with both labels.

All bots run in paper mode. No real money is placed on any exchange. Open positions are repriced every 15 minutes from live market prices, and each paper fill is charged a 0.04% fee on its notional value; the pairs spread template pays that fee on both legs. Rankings are recomputed every hour, and the public table is published with a delay, so what you see is a snapshot of the recent past rather than a live feed. Sharpe ratios are annualized from daily returns using 365 periods a year, and a bot needs at least seven daily returns before its Sharpe means anything.

The league score, piece by piece

Each ranked bot gets a league score from 0 to 100. It is a weighted blend of six sub-scores, each also on a 0–100 scale. Total return counts for 30%, drawdown for 25%, Sharpe for 15%, recent performance for 15%, stability for 10% and data quality for 5%.

The return score is 50 plus 3 times the total return in percent, so a flat bot sits at 50 and every percentage point of gain or loss moves it by 3. The drawdown score is 100 plus 4 times the maximum drawdown in percent; because drawdown is negative, a bot that never fell keeps 100 and a 10% drawdown drops it to 60. The Sharpe score is 50 plus 20 times the Sharpe ratio, so a Sharpe of 2.5 or more already hits the 100 cap. The stability score is 100 minus 5 times the absolute drawdown in percent. Every sub-score is clamped between 0 and 100.

The data-quality score reflects how much the record can be trusted. It falls by 20 points if the equity curve has fewer than 4 points, by 30 if the bot has never synced, by 25 if its data is more than 24 hours stale and by 20 if it is paused. In practice the table publishes this as the data confidence score next to each row.

Notice what the weights reward. Return matters most, but drawdown and stability together carry 35% and both punish deep falls, so a bot that earns a little while barely dipping can outrank one that earns more with large swings. In the 7 October 2026 snapshot, for example, a Regime Switch Bot on ETH/USDT with a total return of −0.013% and a maximum drawdown of −0.017% sat in third place with a score of 77.23, ahead of bots that had made small gains. That is by design: the score is a risk-adjusted ranking, not a profit ranking.

Worked example: scoring one row from the 7 October 2026 snapshot

Take the first-ranked row in the delayed snapshot published on 7 October 2026: a DCA Bot on ETH/USDT, measured from 9 September 2026 to 6 October 2026. It shows a total return of 1.189117%, a maximum drawdown of −0.252680%, a Sharpe of 9.094479, a data confidence score of 95.00 and a league score of 78.544182.

Return score: 50 + 3 × 1.189117 = 53.567. Drawdown score: 100 + 4 × (−0.252680) = 98.989. Sharpe score: 50 + 20 × 9.094479 = 231.890, which is clamped to 100. Stability score: 100 − 5 × 0.252680 = 98.737. Data-quality score: 95, taken from the row's data confidence.

Now apply the weights: 0.30 × 53.567 = 16.070; 0.25 × 98.989 = 24.747; 0.15 × 100 = 15.000; 0.10 × 98.737 = 9.874; 0.05 × 95 = 4.750. The recent-performance component is not shown as its own column: in the current formula version it uses the same calculation as the return score, so it adds 0.15 × 53.567 = 8.035. The six parts add up to 78.476, within 0.07 points of the published 78.544.

Two lessons stand out. First, the Sharpe of 9.09 looks extraordinary, but it comes from under a month of daily returns annualized over 365 days, and the cap means the score treats it exactly like a Sharpe of 2.5. Second, most of this bot's points come from not losing much, not from earning much: the return score alone is only 53.6.

Bots with no trades yet

A bot that has not filled a single order has no return, no drawdown and no meaningful Sharpe. Ranking it would be misleading, because its untouched record would score well on drawdown and stability simply by doing nothing. For that reason bots with zero fills are listed separately under No trades yet and are not ranked. To see why this matters, imagine a perfectly flat record with data quality 80 and a recent-performance score of 50: it would collect 0.30 × 50 + 0.25 × 100 + 0.15 × 50 + 0.15 × 50 + 0.10 × 100 + 0.05 × 80 = 69.0 points without taking any risk at all.

What a leaderboard cannot tell you

Short windows are the first limit. The rows in this snapshot cover roughly four weeks. Crypto markets can stay in one regime for months, so a grid bot that shines in a sideways month or a momentum bot that shines in a trending one may simply be in the right weather. Four weeks is far too short to tell skill from conditions.

Luck is the second. When many bots run at once, some will finish at the top by chance even if none has a real edge. The more entries a table has, the more impressive the best one will look purely through randomness.

Survivorship is the third. A leaderboard only shows the bots that are still on it. If weak bots are stopped, replaced or renamed, the surviving set looks better than the strategies really were. Reading the whole table, including the bottom and the No trades yet list, gives a fairer picture than reading the top three rows.

Finally, paper results are not live results. Paper fills do not move the market, do not wait in a queue and do not suffer outages, so live trading with the same rules will usually do worse.

Reading rankings with risk in mind

Use the leaderboard to learn how different rule families behave under the same conditions, not to pick something to copy. Always check the trading mode, the measurement window, the drawdown and the data confidence before looking at the rank. Nothing in the league is investment advice, past paper results do not predict future returns, and any real trading can lose money, including more than you expect when leverage is involved.

See it in the data

Related guides

Metrics explained with real backtests

Educational content, not investment advice. Guide content and league rankings are independent of any exchange or sponsor relationship.