← All writing

Research · HFM

58 Trading Hypotheses. Most Died. Here's the Framework That Killed Them.

48 of 58 hypotheses were rejected. This is the five-gate kill pipeline that caught them: costs, robustness, null benchmarks, a sealed holdout, and a contamination audit. Each gate is a different question. Run the cheap ones first.

58 hypotheses went in. 48 were rejected, 9 remained promising, and 1 could not be evaluated because the required data was unavailable.

None of the 9 earned the label “edge.” They earned the right to be tested further.

Some ideas should die before a backtest is ever run. The rest should face the cheapest falsification test first. A hypothesis that cannot answer “who pays us?” and “why does this persist?” is not a hypothesis. It is a guess with a Python file attached.

The kill pipeline that processed all 58, in the order most ideas die:

hypothesis idea
    ↓
cost filter        (does it survive fees and turnover at all?)
    ↓
robustness check   (is the result a plateau or a lucky parameter?)
    ↓
null benchmark     (does it beat a structure-preserving random version?)
    ↓
untouched holdout  (does it hold on a year it never saw?)
    ↓
contamination audit (is the result real or did the code lie to me?)
    ↓
survived research → forward / shadow validation → capital
Vertical flowchart of the five-gate hypothesis evaluation pipeline, from input hypothesis through costs, robustness, null benchmark, untouched holdout, and contamination audit, ending at a locked capital node.
The full kill pipeline. Each gate catches a different failure mode. Capital is behind a locked door even after all five pass.

Each gate catches a different type of false alpha. A strategy can clear the first four and die on the fifth. Passing all five is not a green light. It is permission to start collecting forward data.

Gate 1: costs

The most common way a backtest looks good: skip costs. Cross-sectional reversal on 15-minute Bybit bars had a gross Sharpe of 7 to 13, positive in 80%+ of out-of-sample windows. The gross effect was clearly present before costs. At taker fees, rebalancing every bar, trading costs ran 5 to 10 times the size of the raw edge. After costs, zero out of 76 out-of-sample windows remained positive.

I tried entry/exit thresholds, hysteresis, coarser timeframes, slower rebalancing. No plateau anywhere near break-even. Adjacent parameter settings swung from +45% to −60%.

Run the cost test before you believe the signal, not after.

Before and after costs: gross Sharpe 7 to 13 with 80 percent of OOS windows positive versus net Sharpe negative 40 with zero of 76 OOS windows positive.
The same cross-sectional reversal signal. Left: before costs. Right: after. Trading costs ran 5 to 10 times the raw edge.

Gate 2: robustness

Daily momentum had one configuration that returned +97% net at taker fees. The lookback sweep: l=7 gave +33%, l=10 gave −23%, l=14 gave +97%, l=21 gave +35%. No plateau. The +97% was a lucky draw from a jagged surface.

Half the return came from the 2021 bull alone. 2022 was negative. A strategy that only makes money in bull markets and loses in bear is more consistent with market beta than neutral alpha.

One good result on a swept parameter is not evidence. The whole surface needs to hold up.

Line chart of daily momentum net return across four lookback values: l=7 gave plus 33 percent, l=10 gave minus 23, l=14 gave plus 97, l=21 gave plus 35. No plateau.
Momentum lookback sweep. The +97% at l=14 is a spike, not a plateau. Adjacent settings swing from positive to deeply negative.

Gate 3: null benchmark

Funding carry looked strong in development: +48.6% net over 4.5 years, Sharpe 1.17. Before trusting it, I ran a null test: scramble which coin's funding the strategy reads, keep everything else identical. The scrambled version returned +28%.

More than half the return was the structural funding premium — exposure to perpetual market mechanics rather than signal-driven selection. The selection component showed evidence beyond this null in development (empirical permutation p=0.024), but it was much smaller than the headline return implied.

Before believing the result, ask whether a random version of the signal would also make money.

Horizontal bar chart comparing observed funding carry return of plus 48.6 percent to scrambled null of plus 28 percent. The 20.6 percent gap is the selection component, with permutation p equals 0.024.
Funding carry: observed vs scrambled null. More than half the return was structural premium, not signal-driven selection.

Gate 4: untouched holdout

The lockbox is a date range fixed in advance, never touched until a single opening event. No research, no tuning, no plotting inside that range until development is complete.

Funding carry passed development and validation. The lockbox was a 12-month period (2025-09 to 2026-09), sealed from day one. The frozen strategy made −0.5% on it. The premium that paid from 2021 to 2024 was absent in the lockbox year. One year of absence is not a permanent verdict. It is exactly the kind of regime dependence the lockbox is designed to expose early.

Once a lockbox is opened, it is burned. Any further confirmation requires a new, still-unseen date range. Development results are hypotheses, and only sealed, untouched data can confirm them.

Timeline from 2021 to 2026 showing a 4.5-year development period with plus 48.6 percent net return and a sealed lockbox from September 2025 to September 2026 that returned minus 0.5 percent.
The lockbox. Sealed before any research. One opening event. The premium that paid from 2021 to 2024 was absent in the lockbox year.

Gate 5: contamination audit

H56 (BTC funding trend signal) had a development t-stat of 2.6. The 1,248 events had 72-hour holding periods, so observations were heavily overlapping. The Newey-West corrected t-stat fell from 2.6 to 1.07, removing the apparent statistical evidence from that test. A real-looking result from a real-looking test, broken by one statistical assumption.

On a separate project, a football prediction model passed closing-line value (CLV) validation with +152 bps, 95% CI excluded zero. A five-check forensic audit found two contaminated features. Four of five checks failed. The metric that reported the good result was the last place to look for the error.

Left panel shows heavily overlapping 72-hour event windows for 1248 events. Right panel shows t-stat dropping from 2.6 under OLS to 1.07 after Newey-West NW-9 correction.
H56: 1,248 events with 72-hour holding periods are not 1,248 independent observations. One statistical correction removed the apparent evidence.

Where the 58 actually died

Each hypothesis is assigned to the first decisive failure that stopped further evaluation. Some failed more than one check.

Kill gateCountExample
Costs~20Price MR: gross Sharpe 7–13, net negative
Robustness~10Daily momentum: +97% was one lucky lookback
Null benchmark~8Funding carry null still returned +28%
Holdout / fresh OOS~7Funding carry −0.5% on sealed year; agent ensemble −35% OOS
No detectable signal / other rejection~3New-listing cohort: corr≈0, null p=0.94
Total rejected48
Data unavailable1Required data unavailable
Remaining PROMISING9Passed development and validation gates
Benjamini–Hochberg (BH) significant at FDR=0.058
Independent mechanisms4–5roughly 4–5 distinct economic bets

9 survivors does not mean 9 strategies. Several share the same economic mechanism with different entry filters. 8 Benjamini–Hochberg (BH) significant hypotheses at FDR=0.05 collapse to roughly 4–5 independent bets on market structure by economic mechanism.

Funnel diagram showing 58 hypotheses split across five kill gates into 48 rejected, 1 data unavailable, and 9 promising, then narrowing to 8 BH-significant and 4 to 5 independent mechanisms.
Full accounting: 58 in, 48 killed at the first decisive failure, 9 promising, 1 data unavailable. 8 BH survivors reduce to 4–5 independent economic bets.

The uncomfortable failures

Two cases that took longer than they should have.

An LLM-generated factor-residual ensemble showed +76% in development, Sharpe 0.6, tuned over multiple iterations. On a fresh holdout it reversed to −35%. The development result held inside the data it was tuned on. The fresh holdout showed what the agent had actually learned: the development period.

In a funding carry reproduction audit a month after the original publication, the holdout result changed from −0.5% to +1.3%. Same code, same data cache, fingerprints matched. I still do not have a clean explanation. I published the discrepancy table rather than quietly resolving it. I scrutinized the original negative result far less than I would have scrutinized a positive one. Negative results get a lower evidentiary standard by default. They should not.

What survived

A handful of hypotheses passed enough gates to justify forward validation. They are currently shadow-deployed: signals recorded in real time, parameters locked, no capital.

None of them earned the right to be trusted based on the backtest. The backtest is permission to collect more data, not permission to trade.

The specific signals are not published here. Forward validation is ongoing and the gate has not been reached.

Steal this framework

Before trusting any strategy result:

[ ] Realistic costs included from the start, not added after
[ ] Parameter robustness: result holds over adjacent settings (plateau, not spike)
[ ] Null benchmark: structure-preserving random version loses
[ ] Untouched holdout: a date range sealed before any research on that data
[ ] Contamination audit: no look-ahead, no survivorship, no feature leakage
[ ] Forward validation: signals collected in real time, parameters locked

Costs and robustness killed roughly half of the 58 before a holdout was needed. Run the cheap gates first.

Six-item research kill checklist: realistic costs, parameter robustness, null benchmark, untouched holdout, contamination audit, forward validation.
The checklist. Run gates 1 and 2 before believing the signal. Gates 4 and 5 are expensive — clear the cheap ones first.
Research Kill Checklist

One-page version of the framework above. I send new research notes when published.

PK / RESEARCH NOTES

Architecture Review Kit

Rules, annotated diffs, and an audit prompt for catching architecture violations that AI reviewers miss.

The real lesson

A high kill rate is not evidence that the pipeline is good. But a research process that rarely kills its own ideas is one I would distrust.

48 of 58 died. If the kill rate were much lower, I would not trust the process.