Written October 2026, looking back at Systems5 min read
Taking the lock off the hot path
Building a stress harness for our Rust market-data processor, finding a global mutex hiding behind a cache, and what changed when I replaced it. Throughput up a third, p99 down from 139ms to 40ms.
From Print.World · Aug 2025 – Present
At Print.World, the trade path for Solana starts with a Rust service I spend a lot of time in: the DEX processor. It reads a gRPC stream of chain transactions, recognizes swaps across the venues we support, keeps caches of pool state up to date, and publishes what it learns to Kafka and Redis for the rest of the backend. If it falls behind, every price and every quote downstream is a little bit wrong.
In January I wanted to know how much headroom it actually had. Not "it seems fine in prod", but a number I could defend. This is the story of building that number, and of the lock it found.
First, a harness I could trust
The rule I keep coming back to, since my first internship, is that you don't optimize what you can't measure repeatably. So before touching any code I wrote a stress harness: a benchmark that drives the processor's real message handling with synthetic swap traffic as fast as it will go for a fixed duration, and records throughput plus the full latency distribution (p50, p95, p99, p999, max).
A few choices mattered:
- Same code path as production. The harness calls the real handlers, caches, and publishers, not a simplified copy. A benchmark of a toy is a benchmark of a toy.
- Fixed duration, many messages. Each run is 65 seconds and pushes millions of messages, so a few slow outliers can't dominate the result.
- Results committed as JSON. Every run writes a timestamped results file into the repo, so "before" and "after" are artifacts anyone can diff, not numbers in my head.
What the latency shape pointed at
The baseline run did about 132,000 messages per second, with a p50 of 7ms and a p99 of 140ms. The median was fine. The tail was not: a 20x gap between p50 and p99 is the signature of something making threads wait on each other.
Reading through the hot path with that in mind (no profiler, just the latency shape and the code), three things stood out:
- A global
Mutexaround a cache. One of the caches was aHashMapplus aVecDequeused as an LRU, behind a single mutex. Every lookup took the lock, and every cache hit calledretainon the deque to move the key to the front, which is O(n) in the size of the cache. So the most common operation, a hit, was also the most expensive one, and it held a lock everyone needed while doing it. - A
RwLock<HashMap>of subscriber sockets. Reads dominated, but writers still stalled every reader, and the publish path touched it on every message. SeqCststores on the publish path. Shared prices were published with sequentially consistent stores. Readers only needed to see a complete value once it was written, which aReleasestore paired withAcquireloads guarantees. On x86 that's a real difference for stores: aSeqCststore compiles to anXCHG, which acts as a full barrier, while aReleasestore is a plainMOV.
The change
None of the fixes are exotic, which is sort of the point:
- the mutex-guarded cache and the socket map both became
DashMap, a concurrent hash map that shards its keys across many internal locks, so threads only contend when they touch the same shard - the O(n) LRU bookkeeping on every hit went away. The cache holds token decimals, which never change for a given token, so exact recency doesn't matter: when it reaches its configured cap it evicts about a tenth of its entries in one batch, keeping memory bounded without per-hit bookkeeping
- price stores on the publish path moved from
SeqCsttoRelease, withAcquireon the read side
Then I ran the exact same harness again.
- Before
- After
Show data
| metric | Before | After |
|---|---|---|
| p50 | 7ms | 7ms |
| Mean | 12ms | 8ms |
| p95 | 32ms | 15ms |
| p99 | 140ms | 40ms |
| p999 | 251ms | 127ms |
Throughput went from about 132,000 to 176,000 messages per second, a 33% increase on the same hardware, with zero errors in either run.
The number I don't quote
While I was at it I also ran the full in-process pipeline with Kafka publishing turned off, and it hit about 375,000 messages per second. It's tempting to put that next to the 132,000 baseline and claim a 20x improvement. That would be wrong: the baseline had Kafka on, and the comparison mixes two different experiments. The honest use of that number is as a ceiling for what the in-process work can do when the broker is out of the picture, and that's the only way I use it.
What I took from it
- Look at the shape of the distribution, not the average. A healthy median with a terrible p99 tells you what kind of problem you have before you open a profiler.
- The expensive operation is often the common one. A cache hit was doing more work than a miss, under a lock. Nobody designed it that way; it accreted.
- Commit the benchmark. Having before and after as files in the repo made the review conversation short, and it means the next person can rerun it instead of trusting me.
The natural next step is wiring the harness into CI with a loose threshold on p99, so any future change that reintroduces contention shows up in review instead of in a dashboard.