How AvalonQuant re-runs strategy backtests (data, execution, costs, metrics, inclusion rules and known limitations), how the paper-traded league is scored, and where funding-rate data comes from.
The site publishes two different kinds of evidence. Strategy backtests are simulations of fixed rules on about ten years of historical prices. The Live Paper Trading record is a forward paper record: strategy templates trade on live prices with simulated fills, and the record only starts on the day a bot is admitted. A backtest says how a rule would have behaved; a paper record says how it is behaving now, over a much shorter window. Pages label which one you are looking at, and the two are never added together or ranked against each other.
Each published backtest uses public daily price data for US-listed ETFs and stocks. The evaluation window is 2,500 trading days, about ten years (currently October 2016 to October 2026), and each strategy also downloads enough earlier history to warm up its indicators, up to roughly one year. All published strategies use prices adjusted for splits and distributions, so dividends and bond-fund interest are included.
The benchmark on every page is SPY bought and held over exactly the same dates. Calendar-year tables, drawdown dates and market-trend splits are computed from the same daily equity curve as the headline numbers.
Signals use only completed daily bars, so a decision made on one day's close is traded at the next session's open. Strategies rebalance weekly or monthly depending on their rules (Kaufman Efficiency Index waits 21 sessions after its last trade), and most also run daily risk checks between rebalances, such as per-position stops or account-level drawdown breakers that move the portfolio into a Treasury-bill or short-Treasury fund.
Positions are sized as weights of the account, with fractional shares allowed. No strategy borrows on margin. Several hold 2x or 3x leveraged ETFs such as TQQQ, UPRO or SOXL; the leveraged funds' daily-reset decay and fees are captured only because their actual market prices are used. Each strategy page lists its own universe, cadence, risk controls and leverage.
Every trade pays a proportional cost on the value traded. The rate depends on the strategy: 0.10% per traded dollar (Fifty-Two Week High Leaderboard), 0.20% of turnover (Canary-Signaled Tactical ETF), or 0.25% on each buy and each sell (the other six, about 0.50% for a full round trip). These rates are meant to cover commission, spread and slippage together. There is no separate market-impact model, no taxes and no borrowing cost, and fills are assumed at the exact opening price. Strategies that trade often are therefore the ones most exposed to the gap between these assumptions and a real account.
CAGR is the annual rate that compounds the starting value into the ending value over the test period. Maximum drawdown is the largest fall from a previous high to a later low on the daily equity curve. The Sharpe ratio is the mean daily return divided by its standard deviation, annualized, with no risk-free rate subtracted. Return ÷ drawdown divides CAGR by the absolute maximum drawdown (a Calmar-style ratio). Risk tiers group strategies by maximum drawdown: up to 15% lower, up to 25% moderate, up to 35% high, above that very high.
The analysis on each strategy page adds calendar-year returns next to SPY, the three deepest drawdowns with peak, trough and recovery dates and what SPY did over the same stretch, the strategy's return during each decline of more than 10% in SPY, and averages by market trend. Trend labels are assigned after the fact: SPY above a rising 200-day average is an uptrend, below a falling one a downtrend, anything else sideways. They describe the past and are not trading signals.
A strategy is shown only if its current code re-runs, the test covers at least three years, it keeps trading through the period (a curve that stays flat for more than 30% of the days is treated as broken), and it returns a real result. Strategies that fail are left out rather than shown with an inflated annual rate. If a re-run fails for a data reason, such as a missing price on the latest day, the previous re-run stays up until the next successful one.
Methodology reviews can remove a strategy even when its numbers look good. In October 2026 we removed strategies that used same-day information in their signals, ranked from stock lists containing only today's survivors, depended on leveraged funds picked after seeing ten years of results, or stopped working once a realistic 0.1% trading cost was applied.
Selection bias: every universe is a fixed list of ETFs that exist today, many of them well-known 3x funds. Funds that closed are not in the lists, and the lists were chosen with present-day knowledge.
In-sample tuning: several strategies' parameters, such as stop levels and exposure caps, were chosen by comparing alternatives on the same ten-year window that is shown, so the results are not an out-of-sample test.
Short histories and one data source: newer leveraged funds only contribute to recent years, and prices come from a single public data source without cross-checking. Taxes, borrowing and market impact are not modelled. Treat every backtest as evidence about how a rule behaved, not a forecast.
League bots trade in Paper mode only. Every 15 minutes their positions are repriced from live market prices and new targets are filled as simulated trades that pay a 0.04% fee on traded value (pairs-spread bots pay it on both legs); perpetual funding is tracked. Rankings are recomputed hourly and published as delayed snapshots without entry prices, stops or order details. Bots with no fills yet are listed as not ranked so that a flat 0% does not read as a result.
The league score (0 to 100) weights six sub-scores: total return 30%, drawdown 25%, Sharpe 15%, recent performance 15%, stability 10% and data quality 5%. The return sub-score is 50 + 3 × return %, the drawdown sub-score 100 + 4 × max drawdown % (drawdown is negative), the Sharpe sub-score 50 + 20 × Sharpe, and stability 100 − 5 × |max drawdown %|, each limited to 0–100. In the current formula version the recent-performance sub-score uses the same calculation as the return sub-score. Data quality starts from the snapshot confidence and loses points when the curve is too short, the bot has not synced for more than 24 hours, or it is paused. Sharpe in the league is annualized from daily returns over 365 days and is only reported after seven daily returns. The score is a ranking heuristic, not a recommendation.
Funding-rate pages read settled funding for each coin's USDT-margined perpetual from the public Binance and Bybit futures APIs, keep 90 days of history and refresh every 30 minutes. Annualized figures scale the funding actually settled in a 7, 30 or 90-day window to one year; they describe recent holding cost, not a yield. If an exchange cannot be reached, its cells say not available; nothing is estimated in its place.
Calculators run entirely in your browser and show the formula they use. They do not model every exchange rule (for example tiered maintenance margin), so treat their output as a reference and check your exchange before trading.
Backtests and paper results are hypothetical and do not guarantee future results. Nothing on the site is investment advice, a trading signal or an offer to manage money.
Last updated 2026-10-08.