I was testing simple, falsifiable crypto baselines and looking for one result worth a second experiment. Cross-sectional reversal looked almost too clean: a Sharpe of 13.44 before fees and slippage, with positive returns in 67 of 76 reported out-of-sample windows.
At 10 basis points per side, net Sharpe fell to −40.13. Not one reported window remained positive. That is what can happen when a short-lived pattern asks for a new portfolio every 15 minutes.
What I tested
- 15-minute Bybit USDT perpetual bars from January 2021 to the pre-lockbox boundary in September 2025;
- an eight-bar, or two-hour, cross-sectional return signal;
- 90-day formation periods followed by 21-day out-of-sample periods;
- 76 non-overlapping reported OOS windows, rolled every 21 days;
- causal 10% annualized volatility targeting, a 96-bar trailing estimate and a 3× leverage cap;
- a sealed lockbox beginning September 1, 2025, which was not opened for this study.
A window is positive when the sum of its per-bar returns exceeds zero. The cache contains 451 symbol directories, but that is not a reconstructed point-in-time liquid universe. Survivorship and availability bias remain possible. The result is a development diagnostic, not evidence of a deployable alpha.
The deliberately boring signal
At each timestamp, the system measures each available asset's return over the previous eight bars, standardizes returns across the cross-section and reverses the sign. Relative losers receive positive weights; relative winners receive negative weights. The vector is normalized to unit gross exposure.
recent = prices.select(
"ts", (pl.col(symbols) / pl.col(symbols).shift(lookback_bars) - 1.0)
)
signal = long.join(stats, on="ts").with_columns(
signal=pl.when(pl.col("standard_deviation") > 0)
.then(-(pl.col("return") - pl.col("mean")) / pl.col("standard_deviation"))
.otherwise(0.0)
)
weights = signal.join(gross, on="ts").with_columns(
weight=pl.when(pl.col("gross_signal") > 0)
.then(pl.col("signal") / pl.col("gross_signal"))
.otherwise(0.0)
)No model is fitted, and there is no feature search hidden behind the result. A weight computed at time t becomes live at t+1:
def live(column: str) -> pl.Expr:
return pl.col(column).shift(1).fill_null(0.0)Without that shift, the backtest would earn the same closing-bar return used to calculate the signal.
The number that made me stop
The zero-transaction-cost baseline still includes funding PnL. Funding belongs to holding a perpetual position rather than transaction execution, so 13.44 is the Sharpe before fees and slippage—not a pure gross-price-return series.
| Metric | Baseline |
|---|---|
| Annualized Sharpe | 13.44 |
| Positive OOS windows | 67 / 76 (88.2%) |
| Maximum drawdown | −11.62% |
| Summed absolute turnover | 29,451 |
Making the backtest pay its bill
Gross PnL uses execution-shifted weights. Turnover is the sum of absolute changes in live positions. Fees and slippage are fixed basis-point charges on turnover; funding is aligned separately with the long/short sign.
fees = turnover * fee_bps / 10_000
slippage = turnover * slippage_bps / 10_000
net = gross - fees - slippage + fundingThis is a rejection model, not an execution simulator. It does not model spread variation, order-book depth, queue position, partial fills, market impact or adverse selection.
| Metric | 6 bps fee + 4 bps slippage |
|---|---|
| Annualized Sharpe | −40.13 |
| Positive OOS windows | 0 / 76 |
| Cost / gross alpha | 4.00× |
| Total return | approximately −100% |
The useful result is not the theatrical −40.13. It is that the modeled trading bill was four times the gross PnL.
Slowing the strategy down
Under a deliberately pessimistic 15 bps per-side scenario, net Sharpe improved as turnover fell:
| Rebalance | Net Sharpe | Positive windows |
|---|---|---|
| Every 15 minutes | −66.00 | 0 / 76 |
| Every 4 hours | −16.24 | 0 / 76 |
| Every day | −4.45 | 10 / 76 |
The friendliest tested corner used daily rebalancing and 3 bps per side. It reached a Sharpe of −0.01, with 39 of 76 positive windows. Slower trading removed much of the damage, but the signal decayed while it waited.
What failed—and what did not
- Research lead: relative two-hour moves contain a repeatable pre-cost pattern in this sample.
- Failed strategy: continuously resizing the book consumes more than the pattern earns.
- Open engineering problem: reduce turnover without waiting so long that the signal disappears.
The result does not establish that “crypto mean reverts.” It shows a strong pre-transaction-cost pattern in this development sample, and that this naive implementation cannot capture it after simple modeled friction.
Reproduction boundary
The public companion repository contains the article, frozen result snapshot, figures, minimal signal and cost code, deterministic tests, and a manifest that fails CI when a frozen artifact drifts. It does not distribute the market-data cache, so readers can verify the evidence package and mechanics but cannot independently regenerate the historical metrics without sourcing the data.
- The cached-symbol universe is not a point-in-time liquidity universe.
- Trading costs are fixed-bps assumptions, not observed live fills.
- The sealed lockbox was not evaluated; these are backtest results, not live returns.
- The 76 windows do not overlap, but they are not 76 independent market regimes.
SHA-256 fd1f506549b8f43ea33029acb3da760dbc88478f12f0122cf980bc725e6e85a9The result I kept
Turnover belongs inside the hypothesis, not in the cleanup after a backtest looks good. A pre-cost Sharpe is not a strategy. It is a claim before execution.
In this run, almost nothing remained.