← All writing

Research13 min read

Running my own book

How I went from reading about systematic trading to running a six-figure systematic book on crypto perpetual futures, every idea I killed along the way, the oil trade I kept, and why I plan on a much lower Sharpe than my backtest shows.

From My own book · Live since Sep 2026

For the last month most of my evenings have gone into one thing: building and running my own systematic book on Hyperliquid, a perpetual futures exchange that lists crypto and, more recently, stock and commodity perps. It's live, it's six figures, and it trades on its own every hour.

This post is the whole arc: what I read, what I tried, what failed (most of it), and what I actually run. I'm going to show the numbers, including the ugly ones, because the ugly ones are where I learned the most.

45Hypotheses pre-registered in the second wave
15%Drawdown budget the book is built from
9,133Of 9,133 trades reproduced live vs research
1.0Live Sharpe I actually plan on
  1. Early Sept 2026Copy-trading and the first systematic waveTens of thousands of leaderboard rows, millions of fills, and nothing that survived.
  2. Mid SeptAbout 45 pre-registered hypothesesCosts per trade, a kill rule for every idea, and a deflated Sharpe for every result. Slow trend survived.
  3. Sept 22 to 24Tournament, blind test, and idea miningThousands of variants, a red team, and a blind run on 2018 to 2019. The oil spread trade came out of this phase.
  4. Sept 25LiveThe trend core and small event sleeves, with every new idea starting in shadow mode.

The scorecard I use

No single number tells you whether a strategy is good, so every result in this post is scored on the same set of metrics. Each one answers a different question, and a strategy has to look reasonable on all of them.

MetricWhat it measuresWhy I track it
Sharpe ratioAnnual return above the risk-free rate, divided by annual volatility (daily returns, annualized over 365 days)Return per unit of risk; the standard way to compare strategies
Sortino ratioLike Sharpe, but only counts downside volatilityTrend following has big up moves; Sharpe penalizes those, Sortino doesn't
CAGRCompound annual growth rateWhat the money actually does over years, after compounding
Annual volatilityHow much returns swing in a typical yearSets position size: I target a volatility, not a return
Max drawdownLargest fall from a peak before recoveringThe pain number. My whole risk budget is built on it
Calmar ratioCAGR divided by max drawdownReturn per unit of worst-case pain; the one I weigh most
Worst monthThe single worst calendar monthWhat a bad month really feels like
Win rate / hit rateShare of trades or months that make moneyShape of returns, not quality: trend wins under half the time
SkewWhether big moves are mostly up or mostly downTrend has positive skew (rare big wins); carry has negative skew (rare big losses)
TurnoverHow often the book trades its whole sizeDrives costs, which kill most short-horizon ideas
Deflated SharpeThe probability a Sharpe is real, given how many variants I tried, the sample length, and fat tailsTrying enough ideas always produces a lucky one

Where the mindset came from

Before writing any code, I read. The books that shaped how I think about this:

  • Robert Carver, Systematic Trading, which is the most practical book I know on sizing, forecasts, and not overfitting
  • Curtis Faith, Way of the Turtle, and Michael Covel, The Trend Following Bible, for the case that simple trend rules, held with discipline, actually get paid
  • Marcos López de Prado, Advances in Financial Machine Learning, and his papers with David Bailey on the deflated Sharpe ratio and the probability of backtest overfitting
  • Giuseppe Paleologo ("Gappy"), Advanced Portfolio Management and The Elements of Quantitative Investing
  • Moskowitz, Ooi and Pedersen on time-series momentum, and Daniel and Moskowitz on momentum crashes

Two people shaped my attitude more than any single book.

Corey Hoffstein (Newfound Research, and the Flirting with Models podcast) taught me about rebalance timing luck: if your strategy rebalances on Fridays, a big part of your backtest is just which Friday you happened to pick. His answer is to split the book into tranches that rebalance at different times, so no single arbitrary choice can make or break you. More broadly, he taught me to be suspicious of any result that depends on a choice I made without a reason.

Gappy shaped how I think about risk. The idea I took most seriously from him is that you should scale exposure with conditions, not switch it on and off, and that the job of a portfolio manager is mostly to survive long enough for a modest edge to compound. I'm risk-averse by temperament, and his writing gave that instinct a framework: decide your drawdown budget first, size everything from it, and treat a strategy you can't explain as a risk, not an opportunity.

The first wave: everything I believed was wrong

I started where a lot of people start on Hyperliquid: copy-trading. Every account's fills are public, so you can rank the "best" traders and follow them. I pulled about 45,000 leaderboard rows and over four million fills, built 14 statistical gates per wallet, and had three independent skeptics attack every candidate.

None survived. The top wallets turned out to be single lucky longs, market-making bots you can't replicate, hidden losses in other accounts, or deposits that looked like profit. A simulated follower with a one-to-two second lag kept a median of about 18% of the leader's return.

Then I tested the systematic families: momentum, mean reversion, carry, cross-sectional ranks, wallet flow. This chart is the most important thing I learned that week:

The first wave: in-sample vs out-of-sample SharpeWalk-forward, 1,135 out-of-sample days, net of costs. The ideas that looked best on the data they were fit to did the worst afterwards.
  • In-sample
  • Out-of-sample
Show data
ideaIn-sampleOut-of-sample
Funding carry2.33-0.05
Cross-asset / weekend2.06-1.11
Vol-regime momentum1.40-0.48
Cross-sectional momentum1.33-0.53
Follow top wallets1.23-1.42
Time-series momentum1.070.36
Mean reversion0.660.04
Combined system0.390.32

The ideas that looked best on the data they were fit to did the worst afterwards. Funding carry went from an in-sample Sharpe of 2.33 to slightly negative. Smart-money flow went from 1.23 to -1.42. Only plain time-series momentum held on to anything.

The second wave: about 45 hypotheses

So I got more rigorous. Every idea was pre-registered with a kill rule before I looked at its results, every cost (fees, slippage by liquidity class, funding) was charged per trade, and every result was judged with a deflated Sharpe that accounts for how many things I'd tried.

The second wave: out-of-sample Sharpe after costsA selection of roughly 45 hypotheses. Only slow trend cleared 1.0, and it is still well short of a deflation-proof result on its own.
Show data
ideaOut-of-sample Sharpe
Slow crypto trend1.03
Trend ensemble0.77
Cross-sectional momentum0.42
Weekend reversal, stock perps0.93
Commodity seasonality0.18
24h reversal-0.21
Liquidation-cascade reversal-0.32
Calendar seasonality-1.18
Perp basis-1.37
Funding carry, hourly-3.38
Smart-money flow, hourly-3.90

Most of it died. Funding carry, liquidation cascades, smart-money flow, seasonality (zero of 516 calendar hypotheses survived a multiple-testing correction), and every intraday rule I tried on the stock and commodity perps. My own favourite, a Bollinger-band fade, made 17 bps per trade gross and cost 14 bps to trade. Net Sharpe: zero.

The survivor was the least exciting idea on the list: slow trend on crypto, long and short, held until the signal flips. Out-of-sample Sharpe 1.03, and still 1.05 when I re-ran it with 189 delisted coins added back, so it wasn't living off survivorship.

The blind test

A backtest you tuned on 2021 to 2026 data can't tell you much about 2021 to 2026. So I froze the rules and ran them on 2018 and 2019, two years of prices I had never looked at, including one of the worst crypto bear markets on record.

Blind test: the frozen trend book on 2018 to 2019Rules frozen first, then run on two years of prices nobody had looked at. Sharpe 1.28. In 2018 it made about 24% while BTC fell about 73%.
  • Trend book
  • BTC buy and hold
Show data
monthTrend bookBTC buy and hold
2017-12-011.00x1.00x
2018-01-011.04x0.75x
2018-02-011.03x0.75x
2018-03-011.06x0.50x
2018-04-011.16x0.67x
2018-05-011.16x0.55x
2018-06-011.20x0.47x
2018-07-011.17x0.56x
2018-08-011.21x0.51x
2018-09-011.18x0.48x
2018-10-011.15x0.46x
2018-11-011.23x0.29x
2018-12-011.24x0.27x
2019-01-011.21x0.25x
2019-02-011.24x0.28x
2019-03-011.28x0.30x
2019-04-011.22x0.39x
2019-05-011.28x0.62x
2019-06-011.29x0.79x
2019-07-011.30x0.73x
2019-08-011.34x0.70x
2019-09-011.36x0.60x
2019-10-011.33x0.67x
2019-11-011.35x0.55x
2019-12-011.34x0.52x

It held up on every metric, not just Sharpe:

Blind test, 2018 to 2019Value
Sharpe / Sortino1.28 / 1.96
Annual return+15.6%
Annual volatility11.9%
Max drawdown7.5%
Months positive67%
Worst month-4.4%
2018 alone+23.9% while BTC fell 73%

That's the result that convinced me the trend core is real: not because the number is big, but because it showed up on data the rules had never seen.

The oil trade

The one trade in the book I find genuinely elegant is in oil.

Hyperliquid lists perps on WTI and Brent crude. When the CME futures market closes for its daily maintenance hour, those perps can't reference the real futures price, so each one's price oracle falls back to a moving average of its own order book. For that hour, WTI and Brent drift apart on unrelated flow, even though the real spread between them has barely moved. When the futures market reopens, both oracles snap back to the real prices.

So just before the reopen, if the spread has drifted far enough from where it was when the market closed, I fade it: long one leg, short the other, and exit once it snaps back. Trading the spread rather than one leg removes the common oil move, which is exactly what killed the simpler version.

The oil spread trade: cumulative net bps per $1 of spread50 trades, March to September 2026. The pessimistic line assumes a worse fill plus 3 bps of extra cost per side. Note the slope flattening: the edge is decaying, and under pessimistic fills it has earned roughly nothing since mid-May.
  • Modelled fills at touch
  • Pessimistic fills
Show data
dateModelled fills at touchPessimistic fills
2026-03-09172124
2026-03-10225182
2026-03-11253204
2026-03-12269202
2026-03-15487289
2026-03-19511315
2026-03-22641439
2026-03-24703473
2026-03-26715479
2026-03-29770545
2026-03-31780541
2026-04-05786533
2026-04-06790531
2026-04-07824549
2026-04-08836554
2026-04-121,131849
2026-04-191,194918
2026-04-201,239952
2026-04-261,2941,006
2026-04-281,3011,000
2026-04-291,4141,103
2026-05-031,4581,127
2026-05-061,4691,127
2026-05-071,5081,149
2026-05-101,5251,157
2026-05-111,5391,159
2026-05-181,5811,192
2026-05-201,5891,181
2026-05-211,6231,203
2026-05-241,5821,152
2026-05-251,6011,136
2026-06-031,6101,125
2026-06-071,6481,117
2026-06-081,6741,115
2026-06-111,6821,107
2026-06-151,6961,098
2026-07-121,7491,095
2026-07-141,7421,075
2026-07-151,7361,056
2026-07-221,7601,090
2026-07-271,7711,101
2026-08-021,8031,110
2026-08-101,8091,104
2026-08-111,8271,107
2026-08-231,8391,100
2026-08-251,8541,102
2026-08-301,8951,126
2026-09-061,8991,119
2026-09-101,9691,186
2026-09-132,0271,214
Oil spread trade, Mar to Sep 2026Value
Trades50
Win rate94%
Average net per trade+40.5 bps
t-statistic5.0
Edge per trade, Mar-Apr vs May-Sep+67 bps vs +22 bps

Over 50 trades, 94% were winners, at about 40 bps net each. The same rule at six other times of day, when the futures are actually trading, loses money, and so does exiting a few seconds before the snap. That's what convinced me the mechanism is real rather than a pattern I found by looking.

It's also sized small, on purpose. The chart's slope flattens: the edge was about 67 bps per trade in March and April and about 22 since. And it's capacity-bound: it stops working somewhere between $100k and $250k per leg because at that size my fills would become a meaningful share of a thin hour's liquidity. It's a clean mechanism that can't be a big trade, and with the edge decaying I'm comfortable describing it.

What I actually run

The final book, in structure:

  • A crypto trend core across roughly 45 to 50 perps, long and short, decided hourly. A fast and slow moving-average pair, switching to a slower pair when short-term volatility spikes. Positions are sized inversely to downside volatility, and exits only happen on a signal flip. Every trailing stop I tested made it worse, because the trend premium lives in the right tail.
  • A breakout sleeve on BTC and ETH, and a daily book-level volatility target with a partial BTC hedge.
  • A handful of small event sleeves on the stock and commodity perps, including the oil trade, each sized to its capacity.
  • Shadow sleeves that log what they would do without trading, and only get capital once their own live record earns it.
Backtest: the live book's construction, 2021 to 2026Growth of 1.0, net of fees, slippage, impact and funding. This is a backtest on the same history the parameters were chosen on, which is exactly why I plan on a much lower live Sharpe than it shows.
  • Full book
  • Trend core alone
Show data
dateFull bookTrend core alone
2020-12-311.00x1.00x
2021-01-171.13x1.13x
2021-02-141.33x1.33x
2021-03-141.39x1.39x
2021-04-111.39x1.39x
2021-05-091.51x1.51x
2021-06-061.38x1.38x
2021-07-041.39x1.39x
2021-08-011.40x1.40x
2021-08-291.63x1.63x
2021-09-261.68x1.68x
2021-10-241.71x1.71x
2021-11-211.70x1.70x
2021-12-191.78x1.78x
2022-01-161.74x1.74x
2022-02-131.76x1.76x
2022-03-131.78x1.78x
2022-04-101.83x1.83x
2022-05-081.92x1.92x
2022-06-052.00x2.00x
2022-07-032.03x2.03x
2022-07-312.00x2.00x
2022-08-282.08x2.08x
2022-09-252.00x2.00x
2022-10-231.98x1.98x
2022-11-202.08x2.08x
2022-12-182.04x2.04x
2023-01-152.12x2.12x
2023-02-122.21x2.21x
2023-03-122.25x2.25x
2023-04-092.18x2.18x
2023-05-072.15x2.15x
2023-06-042.14x2.14x
2023-07-022.21x2.21x
2023-07-302.13x2.13x
2023-08-272.26x2.26x
2023-09-242.25x2.25x
2023-10-222.24x2.24x
2023-11-192.49x2.49x
2023-12-172.61x2.61x
2024-01-142.68x2.68x
2024-02-112.63x2.63x
2024-03-103.22x3.22x
2024-04-073.10x3.10x
2024-05-053.10x3.10x
2024-06-023.07x3.07x
2024-06-303.15x3.15x
2024-07-283.11x3.11x
2024-08-253.18x3.18x
2024-09-223.11x3.11x
2024-10-203.04x3.04x
2024-11-173.23x3.23x
2024-12-153.59x3.59x
2025-01-123.31x3.31x
2025-02-093.45x3.45x
2025-03-093.54x3.54x
2025-04-063.54x3.54x
2025-05-043.58x3.58x
2025-06-013.69x3.69x
2025-06-293.62x3.62x
2025-07-273.94x3.94x
2025-08-243.76x3.75x
2025-09-213.63x3.61x
2025-10-193.74x3.70x
2025-11-163.89x3.83x
2025-12-143.90x3.83x
2026-01-113.93x3.84x
2026-02-084.18x4.06x
2026-03-084.22x4.03x
2026-04-054.21x3.96x
2026-05-034.31x3.96x
2026-05-314.38x3.94x
2026-06-284.67x4.11x
2026-07-264.60x3.92x
2026-08-235.06x4.29x
2026-09-205.34x4.37x
2026-09-225.40x4.42x

The full scorecard for the live book's construction, split by window. The windows matter: 2021 to 2023 is where the design was chosen, 2024 was the first real check, and 2025 to 2026 is where the most recent choices were made, so it flatters them.

WindowSharpeCAGRVolatilityMax drawdownCalmarWorst month
2021-23 (design)1.7335.3%18.5%9.0%3.93-4.7%
2024 (first check)0.6910.8%17.1%10.8%0.99-6.8%
2025-262.2033.5%13.6%6.6%5.08-4.5%
2021-26 (full)1.6430.1%16.9%12.0%2.52-6.8%

Sharpe here is computed from daily returns and annualized over 365 days, so it won't exactly equal CAGR divided by volatility.

CAGR vs max drawdown, by windowBacktest of the live book's construction. 2024 is the honest year: decent growth, but a drawdown about as large as the return, a Calmar near 1.
  • CAGR
  • Max drawdown
Show data
windowCAGRMax drawdown
2021-23 (design)35%9.0%
2024 (first check)11%11%
2025-2634%6.6%
2021-26 (full)30%12%

The number I keep coming back to is 2024. Sharpe 0.69, Calmar 0.99: a real but unexciting year, and the cleanest out-of-sample read on the core. That's much closer to what I expect live than the full-period figures.

Over the full period that's a Sharpe of 1.64 and a CAGR of about 30%. I plan on a Sharpe of 1.0 to 1.1 live, and I expect the drawdown ladder to fire at some point: at a Sharpe of 1 and about 17% volatility, an unconstrained four-year drawdown is more likely around 20% than 12%. The parameters were chosen on the same history, my coin exclusions were picked while looking at recent data, and haircutting for every trial behind the book takes the 1.64 to roughly 0.8 to 0.9 (the deflated Sharpe probability that the result is real still comes out at 0.97 to 0.98). It also takes something like four years of live data to reliably tell a Sharpe of 1 from 0. I'd rather be pleasantly surprised.

Risk comes first

The book is sized backwards from a 15% drawdown budget, the row I picked from a frontier before choosing any sizes, with 20% as the hard stop. The controls:

  • at 10% below the peak, exposure halves; at 20%, the book goes flat
  • an intraday crash breaker that halves or stops trading on a sharp fall from the 24-hour high
  • reduce-only stops resting on the exchange as a backstop, well before liquidation
  • no new entries above a margin threshold, hard gross caps, and a breaker if the exchange starts rejecting orders
  • stale market data means no new entries, and then closing out
  • any order the system didn't place itself stops the system
The drawdown ladderExposure steps down automatically as the book falls from its peak. No discretion, no override in the moment.

On a replay of the October 2025 crash, these controls ended the day down 5.4%, versus -14.1% without them.

Live versus backtest

The rule I'm strictest about: live trading must reproduce the research exactly. The live engine is a pure, replayable function of time, market data, and targets, so the same code runs in research and in production. Replaying the research history through the live code reproduces 9,133 of 9,133 trend trades, with daily returns matching to floating-point precision.

Every 60 seconds the runner reconciles its view of orders and positions against the exchange, and if they disagree, the exchange wins and that coin freezes until I understand why.

The first live fills also let me check the cost model:

Live slippage vs the cost model, by liquidity class608 live fills over the first three days, measured against the arrival mid, fees excluded. Small sample, but the model was conservative everywhere.
  • Modelled (bps)
  • Live (bps)
Show data
clsModelled (bps)Live (bps)
all core6.801.12
majors (BTC, ETH)1.810.15
liquid alts5.781.76
other alts12.981.53

Across 608 live fills, slippage came in at about 1.1 bps against the arrival mid, versus 6.8 bps modelled. It's a small sample, but it means the backtest was charging itself conservatively, which is the direction you want to be wrong in.

What I took from it

Most of the value of this project was in the things I stopped believing. Copy-trading the best wallets doesn't work. High in-sample Sharpe is a warning sign, not a credential. Costs decide most intraday ideas before the signal does. And the thing that survived every test I could think of was the most boring idea on the list, held with discipline and sized so that being wrong is survivable.

That's also the work I want to keep doing: careful research, honest about luck, with an execution path I can trust to do exactly what the research says.