Research13 min read
Running my own book
How I went from reading about systematic trading to running a six-figure systematic book on crypto perpetual futures, every idea I killed along the way, the oil trade I kept, and why I plan on a much lower Sharpe than my backtest shows.
From My own book · Live since Sep 2026
For the last month most of my evenings have gone into one thing: building and running my own systematic book on Hyperliquid, a perpetual futures exchange that lists crypto and, more recently, stock and commodity perps. It's live, it's six figures, and it trades on its own every hour.
This post is the whole arc: what I read, what I tried, what failed (most of it), and what I actually run. I'm going to show the numbers, including the ugly ones, because the ugly ones are where I learned the most.
- Early Sept 2026Copy-trading and the first systematic waveTens of thousands of leaderboard rows, millions of fills, and nothing that survived.
- Mid SeptAbout 45 pre-registered hypothesesCosts per trade, a kill rule for every idea, and a deflated Sharpe for every result. Slow trend survived.
- Sept 22 to 24Tournament, blind test, and idea miningThousands of variants, a red team, and a blind run on 2018 to 2019. The oil spread trade came out of this phase.
- Sept 25LiveThe trend core and small event sleeves, with every new idea starting in shadow mode.
The scorecard I use
No single number tells you whether a strategy is good, so every result in this post is scored on the same set of metrics. Each one answers a different question, and a strategy has to look reasonable on all of them.
| Metric | What it measures | Why I track it |
|---|---|---|
| Sharpe ratio | Annual return above the risk-free rate, divided by annual volatility (daily returns, annualized over 365 days) | Return per unit of risk; the standard way to compare strategies |
| Sortino ratio | Like Sharpe, but only counts downside volatility | Trend following has big up moves; Sharpe penalizes those, Sortino doesn't |
| CAGR | Compound annual growth rate | What the money actually does over years, after compounding |
| Annual volatility | How much returns swing in a typical year | Sets position size: I target a volatility, not a return |
| Max drawdown | Largest fall from a peak before recovering | The pain number. My whole risk budget is built on it |
| Calmar ratio | CAGR divided by max drawdown | Return per unit of worst-case pain; the one I weigh most |
| Worst month | The single worst calendar month | What a bad month really feels like |
| Win rate / hit rate | Share of trades or months that make money | Shape of returns, not quality: trend wins under half the time |
| Skew | Whether big moves are mostly up or mostly down | Trend has positive skew (rare big wins); carry has negative skew (rare big losses) |
| Turnover | How often the book trades its whole size | Drives costs, which kill most short-horizon ideas |
| Deflated Sharpe | The probability a Sharpe is real, given how many variants I tried, the sample length, and fat tails | Trying enough ideas always produces a lucky one |
Where the mindset came from
Before writing any code, I read. The books that shaped how I think about this:
- Robert Carver, Systematic Trading, which is the most practical book I know on sizing, forecasts, and not overfitting
- Curtis Faith, Way of the Turtle, and Michael Covel, The Trend Following Bible, for the case that simple trend rules, held with discipline, actually get paid
- Marcos López de Prado, Advances in Financial Machine Learning, and his papers with David Bailey on the deflated Sharpe ratio and the probability of backtest overfitting
- Giuseppe Paleologo ("Gappy"), Advanced Portfolio Management and The Elements of Quantitative Investing
- Moskowitz, Ooi and Pedersen on time-series momentum, and Daniel and Moskowitz on momentum crashes
Two people shaped my attitude more than any single book.
Corey Hoffstein (Newfound Research, and the Flirting with Models podcast) taught me about rebalance timing luck: if your strategy rebalances on Fridays, a big part of your backtest is just which Friday you happened to pick. His answer is to split the book into tranches that rebalance at different times, so no single arbitrary choice can make or break you. More broadly, he taught me to be suspicious of any result that depends on a choice I made without a reason.
Gappy shaped how I think about risk. The idea I took most seriously from him is that you should scale exposure with conditions, not switch it on and off, and that the job of a portfolio manager is mostly to survive long enough for a modest edge to compound. I'm risk-averse by temperament, and his writing gave that instinct a framework: decide your drawdown budget first, size everything from it, and treat a strategy you can't explain as a risk, not an opportunity.
The first wave: everything I believed was wrong
I started where a lot of people start on Hyperliquid: copy-trading. Every account's fills are public, so you can rank the "best" traders and follow them. I pulled about 45,000 leaderboard rows and over four million fills, built 14 statistical gates per wallet, and had three independent skeptics attack every candidate.
None survived. The top wallets turned out to be single lucky longs, market-making bots you can't replicate, hidden losses in other accounts, or deposits that looked like profit. A simulated follower with a one-to-two second lag kept a median of about 18% of the leader's return.
Then I tested the systematic families: momentum, mean reversion, carry, cross-sectional ranks, wallet flow. This chart is the most important thing I learned that week:
- In-sample
- Out-of-sample
Show data
| idea | In-sample | Out-of-sample |
|---|---|---|
| Funding carry | 2.33 | -0.05 |
| Cross-asset / weekend | 2.06 | -1.11 |
| Vol-regime momentum | 1.40 | -0.48 |
| Cross-sectional momentum | 1.33 | -0.53 |
| Follow top wallets | 1.23 | -1.42 |
| Time-series momentum | 1.07 | 0.36 |
| Mean reversion | 0.66 | 0.04 |
| Combined system | 0.39 | 0.32 |
The ideas that looked best on the data they were fit to did the worst afterwards. Funding carry went from an in-sample Sharpe of 2.33 to slightly negative. Smart-money flow went from 1.23 to -1.42. Only plain time-series momentum held on to anything.
The second wave: about 45 hypotheses
So I got more rigorous. Every idea was pre-registered with a kill rule before I looked at its results, every cost (fees, slippage by liquidity class, funding) was charged per trade, and every result was judged with a deflated Sharpe that accounts for how many things I'd tried.
Show data
| idea | Out-of-sample Sharpe |
|---|---|
| Slow crypto trend | 1.03 |
| Trend ensemble | 0.77 |
| Cross-sectional momentum | 0.42 |
| Weekend reversal, stock perps | 0.93 |
| Commodity seasonality | 0.18 |
| 24h reversal | -0.21 |
| Liquidation-cascade reversal | -0.32 |
| Calendar seasonality | -1.18 |
| Perp basis | -1.37 |
| Funding carry, hourly | -3.38 |
| Smart-money flow, hourly | -3.90 |
Most of it died. Funding carry, liquidation cascades, smart-money flow, seasonality (zero of 516 calendar hypotheses survived a multiple-testing correction), and every intraday rule I tried on the stock and commodity perps. My own favourite, a Bollinger-band fade, made 17 bps per trade gross and cost 14 bps to trade. Net Sharpe: zero.
The survivor was the least exciting idea on the list: slow trend on crypto, long and short, held until the signal flips. Out-of-sample Sharpe 1.03, and still 1.05 when I re-ran it with 189 delisted coins added back, so it wasn't living off survivorship.
The blind test
A backtest you tuned on 2021 to 2026 data can't tell you much about 2021 to 2026. So I froze the rules and ran them on 2018 and 2019, two years of prices I had never looked at, including one of the worst crypto bear markets on record.
- Trend book
- BTC buy and hold
Show data
| month | Trend book | BTC buy and hold |
|---|---|---|
| 2017-12-01 | 1.00x | 1.00x |
| 2018-01-01 | 1.04x | 0.75x |
| 2018-02-01 | 1.03x | 0.75x |
| 2018-03-01 | 1.06x | 0.50x |
| 2018-04-01 | 1.16x | 0.67x |
| 2018-05-01 | 1.16x | 0.55x |
| 2018-06-01 | 1.20x | 0.47x |
| 2018-07-01 | 1.17x | 0.56x |
| 2018-08-01 | 1.21x | 0.51x |
| 2018-09-01 | 1.18x | 0.48x |
| 2018-10-01 | 1.15x | 0.46x |
| 2018-11-01 | 1.23x | 0.29x |
| 2018-12-01 | 1.24x | 0.27x |
| 2019-01-01 | 1.21x | 0.25x |
| 2019-02-01 | 1.24x | 0.28x |
| 2019-03-01 | 1.28x | 0.30x |
| 2019-04-01 | 1.22x | 0.39x |
| 2019-05-01 | 1.28x | 0.62x |
| 2019-06-01 | 1.29x | 0.79x |
| 2019-07-01 | 1.30x | 0.73x |
| 2019-08-01 | 1.34x | 0.70x |
| 2019-09-01 | 1.36x | 0.60x |
| 2019-10-01 | 1.33x | 0.67x |
| 2019-11-01 | 1.35x | 0.55x |
| 2019-12-01 | 1.34x | 0.52x |
It held up on every metric, not just Sharpe:
| Blind test, 2018 to 2019 | Value |
|---|---|
| Sharpe / Sortino | 1.28 / 1.96 |
| Annual return | +15.6% |
| Annual volatility | 11.9% |
| Max drawdown | 7.5% |
| Months positive | 67% |
| Worst month | -4.4% |
| 2018 alone | +23.9% while BTC fell 73% |
That's the result that convinced me the trend core is real: not because the number is big, but because it showed up on data the rules had never seen.
The oil trade
The one trade in the book I find genuinely elegant is in oil.
Hyperliquid lists perps on WTI and Brent crude. When the CME futures market closes for its daily maintenance hour, those perps can't reference the real futures price, so each one's price oracle falls back to a moving average of its own order book. For that hour, WTI and Brent drift apart on unrelated flow, even though the real spread between them has barely moved. When the futures market reopens, both oracles snap back to the real prices.
So just before the reopen, if the spread has drifted far enough from where it was when the market closed, I fade it: long one leg, short the other, and exit once it snaps back. Trading the spread rather than one leg removes the common oil move, which is exactly what killed the simpler version.
- Modelled fills at touch
- Pessimistic fills
Show data
| date | Modelled fills at touch | Pessimistic fills |
|---|---|---|
| 2026-03-09 | 172 | 124 |
| 2026-03-10 | 225 | 182 |
| 2026-03-11 | 253 | 204 |
| 2026-03-12 | 269 | 202 |
| 2026-03-15 | 487 | 289 |
| 2026-03-19 | 511 | 315 |
| 2026-03-22 | 641 | 439 |
| 2026-03-24 | 703 | 473 |
| 2026-03-26 | 715 | 479 |
| 2026-03-29 | 770 | 545 |
| 2026-03-31 | 780 | 541 |
| 2026-04-05 | 786 | 533 |
| 2026-04-06 | 790 | 531 |
| 2026-04-07 | 824 | 549 |
| 2026-04-08 | 836 | 554 |
| 2026-04-12 | 1,131 | 849 |
| 2026-04-19 | 1,194 | 918 |
| 2026-04-20 | 1,239 | 952 |
| 2026-04-26 | 1,294 | 1,006 |
| 2026-04-28 | 1,301 | 1,000 |
| 2026-04-29 | 1,414 | 1,103 |
| 2026-05-03 | 1,458 | 1,127 |
| 2026-05-06 | 1,469 | 1,127 |
| 2026-05-07 | 1,508 | 1,149 |
| 2026-05-10 | 1,525 | 1,157 |
| 2026-05-11 | 1,539 | 1,159 |
| 2026-05-18 | 1,581 | 1,192 |
| 2026-05-20 | 1,589 | 1,181 |
| 2026-05-21 | 1,623 | 1,203 |
| 2026-05-24 | 1,582 | 1,152 |
| 2026-05-25 | 1,601 | 1,136 |
| 2026-06-03 | 1,610 | 1,125 |
| 2026-06-07 | 1,648 | 1,117 |
| 2026-06-08 | 1,674 | 1,115 |
| 2026-06-11 | 1,682 | 1,107 |
| 2026-06-15 | 1,696 | 1,098 |
| 2026-07-12 | 1,749 | 1,095 |
| 2026-07-14 | 1,742 | 1,075 |
| 2026-07-15 | 1,736 | 1,056 |
| 2026-07-22 | 1,760 | 1,090 |
| 2026-07-27 | 1,771 | 1,101 |
| 2026-08-02 | 1,803 | 1,110 |
| 2026-08-10 | 1,809 | 1,104 |
| 2026-08-11 | 1,827 | 1,107 |
| 2026-08-23 | 1,839 | 1,100 |
| 2026-08-25 | 1,854 | 1,102 |
| 2026-08-30 | 1,895 | 1,126 |
| 2026-09-06 | 1,899 | 1,119 |
| 2026-09-10 | 1,969 | 1,186 |
| 2026-09-13 | 2,027 | 1,214 |
| Oil spread trade, Mar to Sep 2026 | Value |
|---|---|
| Trades | 50 |
| Win rate | 94% |
| Average net per trade | +40.5 bps |
| t-statistic | 5.0 |
| Edge per trade, Mar-Apr vs May-Sep | +67 bps vs +22 bps |
Over 50 trades, 94% were winners, at about 40 bps net each. The same rule at six other times of day, when the futures are actually trading, loses money, and so does exiting a few seconds before the snap. That's what convinced me the mechanism is real rather than a pattern I found by looking.
It's also sized small, on purpose. The chart's slope flattens: the edge was about 67 bps per trade in March and April and about 22 since. And it's capacity-bound: it stops working somewhere between $100k and $250k per leg because at that size my fills would become a meaningful share of a thin hour's liquidity. It's a clean mechanism that can't be a big trade, and with the edge decaying I'm comfortable describing it.
What I actually run
The final book, in structure:
- A crypto trend core across roughly 45 to 50 perps, long and short, decided hourly. A fast and slow moving-average pair, switching to a slower pair when short-term volatility spikes. Positions are sized inversely to downside volatility, and exits only happen on a signal flip. Every trailing stop I tested made it worse, because the trend premium lives in the right tail.
- A breakout sleeve on BTC and ETH, and a daily book-level volatility target with a partial BTC hedge.
- A handful of small event sleeves on the stock and commodity perps, including the oil trade, each sized to its capacity.
- Shadow sleeves that log what they would do without trading, and only get capital once their own live record earns it.
- Full book
- Trend core alone
Show data
| date | Full book | Trend core alone |
|---|---|---|
| 2020-12-31 | 1.00x | 1.00x |
| 2021-01-17 | 1.13x | 1.13x |
| 2021-02-14 | 1.33x | 1.33x |
| 2021-03-14 | 1.39x | 1.39x |
| 2021-04-11 | 1.39x | 1.39x |
| 2021-05-09 | 1.51x | 1.51x |
| 2021-06-06 | 1.38x | 1.38x |
| 2021-07-04 | 1.39x | 1.39x |
| 2021-08-01 | 1.40x | 1.40x |
| 2021-08-29 | 1.63x | 1.63x |
| 2021-09-26 | 1.68x | 1.68x |
| 2021-10-24 | 1.71x | 1.71x |
| 2021-11-21 | 1.70x | 1.70x |
| 2021-12-19 | 1.78x | 1.78x |
| 2022-01-16 | 1.74x | 1.74x |
| 2022-02-13 | 1.76x | 1.76x |
| 2022-03-13 | 1.78x | 1.78x |
| 2022-04-10 | 1.83x | 1.83x |
| 2022-05-08 | 1.92x | 1.92x |
| 2022-06-05 | 2.00x | 2.00x |
| 2022-07-03 | 2.03x | 2.03x |
| 2022-07-31 | 2.00x | 2.00x |
| 2022-08-28 | 2.08x | 2.08x |
| 2022-09-25 | 2.00x | 2.00x |
| 2022-10-23 | 1.98x | 1.98x |
| 2022-11-20 | 2.08x | 2.08x |
| 2022-12-18 | 2.04x | 2.04x |
| 2023-01-15 | 2.12x | 2.12x |
| 2023-02-12 | 2.21x | 2.21x |
| 2023-03-12 | 2.25x | 2.25x |
| 2023-04-09 | 2.18x | 2.18x |
| 2023-05-07 | 2.15x | 2.15x |
| 2023-06-04 | 2.14x | 2.14x |
| 2023-07-02 | 2.21x | 2.21x |
| 2023-07-30 | 2.13x | 2.13x |
| 2023-08-27 | 2.26x | 2.26x |
| 2023-09-24 | 2.25x | 2.25x |
| 2023-10-22 | 2.24x | 2.24x |
| 2023-11-19 | 2.49x | 2.49x |
| 2023-12-17 | 2.61x | 2.61x |
| 2024-01-14 | 2.68x | 2.68x |
| 2024-02-11 | 2.63x | 2.63x |
| 2024-03-10 | 3.22x | 3.22x |
| 2024-04-07 | 3.10x | 3.10x |
| 2024-05-05 | 3.10x | 3.10x |
| 2024-06-02 | 3.07x | 3.07x |
| 2024-06-30 | 3.15x | 3.15x |
| 2024-07-28 | 3.11x | 3.11x |
| 2024-08-25 | 3.18x | 3.18x |
| 2024-09-22 | 3.11x | 3.11x |
| 2024-10-20 | 3.04x | 3.04x |
| 2024-11-17 | 3.23x | 3.23x |
| 2024-12-15 | 3.59x | 3.59x |
| 2025-01-12 | 3.31x | 3.31x |
| 2025-02-09 | 3.45x | 3.45x |
| 2025-03-09 | 3.54x | 3.54x |
| 2025-04-06 | 3.54x | 3.54x |
| 2025-05-04 | 3.58x | 3.58x |
| 2025-06-01 | 3.69x | 3.69x |
| 2025-06-29 | 3.62x | 3.62x |
| 2025-07-27 | 3.94x | 3.94x |
| 2025-08-24 | 3.76x | 3.75x |
| 2025-09-21 | 3.63x | 3.61x |
| 2025-10-19 | 3.74x | 3.70x |
| 2025-11-16 | 3.89x | 3.83x |
| 2025-12-14 | 3.90x | 3.83x |
| 2026-01-11 | 3.93x | 3.84x |
| 2026-02-08 | 4.18x | 4.06x |
| 2026-03-08 | 4.22x | 4.03x |
| 2026-04-05 | 4.21x | 3.96x |
| 2026-05-03 | 4.31x | 3.96x |
| 2026-05-31 | 4.38x | 3.94x |
| 2026-06-28 | 4.67x | 4.11x |
| 2026-07-26 | 4.60x | 3.92x |
| 2026-08-23 | 5.06x | 4.29x |
| 2026-09-20 | 5.34x | 4.37x |
| 2026-09-22 | 5.40x | 4.42x |
The full scorecard for the live book's construction, split by window. The windows matter: 2021 to 2023 is where the design was chosen, 2024 was the first real check, and 2025 to 2026 is where the most recent choices were made, so it flatters them.
| Window | Sharpe | CAGR | Volatility | Max drawdown | Calmar | Worst month |
|---|---|---|---|---|---|---|
| 2021-23 (design) | 1.73 | 35.3% | 18.5% | 9.0% | 3.93 | -4.7% |
| 2024 (first check) | 0.69 | 10.8% | 17.1% | 10.8% | 0.99 | -6.8% |
| 2025-26 | 2.20 | 33.5% | 13.6% | 6.6% | 5.08 | -4.5% |
| 2021-26 (full) | 1.64 | 30.1% | 16.9% | 12.0% | 2.52 | -6.8% |
Sharpe here is computed from daily returns and annualized over 365 days, so it won't exactly equal CAGR divided by volatility.
- CAGR
- Max drawdown
Show data
| window | CAGR | Max drawdown |
|---|---|---|
| 2021-23 (design) | 35% | 9.0% |
| 2024 (first check) | 11% | 11% |
| 2025-26 | 34% | 6.6% |
| 2021-26 (full) | 30% | 12% |
The number I keep coming back to is 2024. Sharpe 0.69, Calmar 0.99: a real but unexciting year, and the cleanest out-of-sample read on the core. That's much closer to what I expect live than the full-period figures.
Over the full period that's a Sharpe of 1.64 and a CAGR of about 30%. I plan on a Sharpe of 1.0 to 1.1 live, and I expect the drawdown ladder to fire at some point: at a Sharpe of 1 and about 17% volatility, an unconstrained four-year drawdown is more likely around 20% than 12%. The parameters were chosen on the same history, my coin exclusions were picked while looking at recent data, and haircutting for every trial behind the book takes the 1.64 to roughly 0.8 to 0.9 (the deflated Sharpe probability that the result is real still comes out at 0.97 to 0.98). It also takes something like four years of live data to reliably tell a Sharpe of 1 from 0. I'd rather be pleasantly surprised.
Risk comes first
The book is sized backwards from a 15% drawdown budget, the row I picked from a frontier before choosing any sizes, with 20% as the hard stop. The controls:
- at 10% below the peak, exposure halves; at 20%, the book goes flat
- an intraday crash breaker that halves or stops trading on a sharp fall from the 24-hour high
- reduce-only stops resting on the exchange as a backstop, well before liquidation
- no new entries above a margin threshold, hard gross caps, and a breaker if the exchange starts rejecting orders
- stale market data means no new entries, and then closing out
- any order the system didn't place itself stops the system
On a replay of the October 2025 crash, these controls ended the day down 5.4%, versus -14.1% without them.
Live versus backtest
The rule I'm strictest about: live trading must reproduce the research exactly. The live engine is a pure, replayable function of time, market data, and targets, so the same code runs in research and in production. Replaying the research history through the live code reproduces 9,133 of 9,133 trend trades, with daily returns matching to floating-point precision.
Every 60 seconds the runner reconciles its view of orders and positions against the exchange, and if they disagree, the exchange wins and that coin freezes until I understand why.
The first live fills also let me check the cost model:
- Modelled (bps)
- Live (bps)
Show data
| cls | Modelled (bps) | Live (bps) |
|---|---|---|
| all core | 6.80 | 1.12 |
| majors (BTC, ETH) | 1.81 | 0.15 |
| liquid alts | 5.78 | 1.76 |
| other alts | 12.98 | 1.53 |
Across 608 live fills, slippage came in at about 1.1 bps against the arrival mid, versus 6.8 bps modelled. It's a small sample, but it means the backtest was charging itself conservatively, which is the direction you want to be wrong in.
What I took from it
Most of the value of this project was in the things I stopped believing. Copy-trading the best wallets doesn't work. High in-sample Sharpe is a warning sign, not a credential. Costs decide most intraday ideas before the signal does. And the thing that survived every test I could think of was the most boring idea on the list, held with discipline and sized so that being wrong is survivable.
That's also the work I want to keep doing: careful research, honest about luck, with an execution path I can trust to do exactly what the research says.