← All writing

Written October 2026, looking back at Systems5 min read

Taking the lock off the hot path

Building a stress harness for our Rust market-data processor, finding a global mutex hiding behind a cache, and what changed when I replaced it. Throughput up a third, p99 down from 139ms to 40ms.

From Print.World · Aug 2025 – Present

At Print.World, the trade path for Solana starts with a Rust service I spend a lot of time in: the DEX processor. It reads a gRPC stream of chain transactions, recognizes swaps across the venues we support, keeps caches of pool state up to date, and publishes what it learns to Kafka and Redis for the rest of the backend. If it falls behind, every price and every quote downstream is a little bit wrong.

In January I wanted to know how much headroom it actually had. Not "it seems fine in prod", but a number I could defend. This is the story of building that number, and of the lock it found.

+33%Throughput, same hardware
139→40p99 latency, milliseconds
31→15p95 latency, milliseconds
0Errors across both 65 second runs

First, a harness I could trust

The rule I keep coming back to, since my first internship, is that you don't optimize what you can't measure repeatably. So before touching any code I wrote a stress harness: a benchmark that drives the processor's real message handling with synthetic swap traffic as fast as it will go for a fixed duration, and records throughput plus the full latency distribution (p50, p95, p99, p999, max).

A few choices mattered:

  • Same code path as production. The harness calls the real handlers, caches, and publishers, not a simplified copy. A benchmark of a toy is a benchmark of a toy.
  • Fixed duration, many messages. Each run is 65 seconds and pushes millions of messages, so a few slow outliers can't dominate the result.
  • Results committed as JSON. Every run writes a timestamped results file into the repo, so "before" and "after" are artifacts anyone can diff, not numbers in my head.

What the latency shape pointed at

The baseline run did about 132,000 messages per second, with a p50 of 7ms and a p99 of 140ms. The median was fine. The tail was not: a 20x gap between p50 and p99 is the signature of something making threads wait on each other.

Reading through the hot path with that in mind (no profiler, just the latency shape and the code), three things stood out:

  1. A global Mutex around a cache. One of the caches was a HashMap plus a VecDeque used as an LRU, behind a single mutex. Every lookup took the lock, and every cache hit called retain on the deque to move the key to the front, which is O(n) in the size of the cache. So the most common operation, a hit, was also the most expensive one, and it held a lock everyone needed while doing it.
  2. A RwLock<HashMap> of subscriber sockets. Reads dominated, but writers still stalled every reader, and the publish path touched it on every message.
  3. SeqCst stores on the publish path. Shared prices were published with sequentially consistent stores. Readers only needed to see a complete value once it was written, which a Release store paired with Acquire loads guarantees. On x86 that's a real difference for stores: a SeqCst store compiles to an XCHG, which acts as a full barrier, while a Release store is a plain MOV.
Before and after on the hot pathEvery message used to queue on one lock, and a cache hit did O(n) work while holding it. After the change, threads only contend when they touch the same shard.

The change

None of the fixes are exotic, which is sort of the point:

  • the mutex-guarded cache and the socket map both became DashMap, a concurrent hash map that shards its keys across many internal locks, so threads only contend when they touch the same shard
  • the O(n) LRU bookkeeping on every hit went away. The cache holds token decimals, which never change for a given token, so exact recency doesn't matter: when it reaches its configured cap it evicts about a tenth of its entries in one batch, keeping memory bounded without per-hit bookkeeping
  • price stores on the publish path moved from SeqCst to Release, with Acquire on the read side

Then I ran the exact same harness again.

Latency before and after (ms)Same harness, same 65 second duration, same machine. The median barely moved; the tail collapsed, which is what you expect when you remove contention rather than make the work itself faster.
  • Before
  • After
Show data
metricBeforeAfter
p507ms7ms
Mean12ms8ms
p9532ms15ms
p99140ms40ms
p999251ms127ms

Throughput went from about 132,000 to 176,000 messages per second, a 33% increase on the same hardware, with zero errors in either run.

The number I don't quote

While I was at it I also ran the full in-process pipeline with Kafka publishing turned off, and it hit about 375,000 messages per second. It's tempting to put that next to the 132,000 baseline and claim a 20x improvement. That would be wrong: the baseline had Kafka on, and the comparison mixes two different experiments. The honest use of that number is as a ceiling for what the in-process work can do when the broker is out of the picture, and that's the only way I use it.

What I took from it

  • Look at the shape of the distribution, not the average. A healthy median with a terrible p99 tells you what kind of problem you have before you open a profiler.
  • The expensive operation is often the common one. A cache hit was doing more work than a miss, under a lock. Nobody designed it that way; it accreted.
  • Commit the benchmark. Having before and after as files in the repo made the review conversation short, and it means the next person can rerun it instead of trusting me.

The natural next step is wiring the harness into CI with a loose threshold on p99, so any future change that reintroduces contention shows up in review instead of in a dashboard.