Systems4 min read
Fix our websocket server, or adopt one?
Our in-house websocket layer was struggling with live perps order books. The bounded fixes looked cheap, adopting Centrifugo looked expensive, and the decision came down to which one I'd still trust in six months.
From Print.World · Aug 2025 – Present
When we started streaming live perpetual futures order books to the terminal, our websocket layer started to creak. This post is about the decision I had to make there: patch the server we already had, or put a dedicated real-time server next to it. It's the kind of decision that doesn't have a right answer, only a defensible one, so I want to show the reasoning, not just the outcome.
The problem
Print.World's terminal gets live data over an in-house websocket service. It had served us well for prices and trades. Order books are different: they update many times a second, every viewer of a market needs them, and a client that misses an update has a wrong book until it gets a fresh snapshot.
Looking at how the existing server handled that, a few structural issues stood out:
- Every instance received every message. Fan-out ran over a single Redis pub/sub channel, so each server instance processed traffic for markets none of its clients were watching.
- Work per socket, not per message. Each outgoing message was serialized and compressed once per connected socket, so the cost scaled with viewers instead of with updates.
- No backpressure and no recovery. A slow client just accumulated a growing send buffer, and a client that reconnected had no way to ask for what it missed.
We'd already had a warning shot: during one incident, pub/sub ingress jumped from about 13 to 60 MB/s and Redis closed the subscriber connection. Perps books would make that traffic heavier, not lighter.
Option one: four bounded fixes
The cheap-looking option was to fix the existing server in place. I wrote down exactly what that would take:
- per-channel subscriptions, so instances only receive markets their clients want
- compress once per message, not once per socket
- a ceiling on each socket's send buffer, so a slow client gets dropped instead of hurting everyone
- sequenced deltas, so clients can detect a gap and recover
Each of those is a reasonable change. Together, they touch every real-time feature on the platform, not just perps. And when I listed them out, I noticed something uncomfortable: I was describing, piece by piece, a subset of what a dedicated real-time server already does.
Option two: Centrifugo, next to what we have
Centrifugo is an open-source real-time messaging server. It already does channel subscriptions, history and recovery with offsets and epochs, and delta compression: sending a patch against the last message instead of the whole thing.
The costs were real, too. There's no license fee, but it's a second service to deploy, monitor, and understand, and the team would have two real-time paths for a while.
| Fix in place | Adopt Centrifugo for perps | |
|---|---|---|
| Blast radius | Every real-time feature | Perps namespace only |
| Recovery after reconnect | Build it | Built in (history, offsets, epochs) |
| Delta compression | Build it | Built in |
| New operational surface | None | One more service |
| Risk of subtle regressions | High, shared code | Low, isolated |
The deciding factor wasn't performance. It was blast radius. Fixing the shared server meant risking regressions in features that had nothing to do with perps, in code every real-time message flows through. Adding a second server meant more to run, but a failure in it could only hurt the thing I was building.
Getting the details right
A few pieces I'm happiest with:
- Gaps are errors, not warnings. Venue order-book feeds are sequenced. If a delta's sequence number isn't exactly the last one plus one, the book is marked stale and the client re-fetches a full snapshot. Applying an update past a gap would silently serve a wrong book, which is worse than briefly serving none.
- One subscription per channel per client. On the terminal side, a reference-counted client means ten components watching the same market share one subscription instead of opening ten.
- Delta compression on the wire. Clients ask for deltas, so after the first full book, each message is a small patch.
Show data
| mode | KB/s |
|---|---|
| Full book every update | 132 |
| Delta compression | 6.40 |
That's about a 20x reduction in bandwidth for the busiest book, measured locally. I'm careful to say "locally": production rollout was still pending its infrastructure when I wrote the design doc, and the doc says so.
What I took from it
List the fixes before you choose them. Writing the four bounded fixes out explicitly is what showed me I'd be rebuilding a real-time server badly. Without the list, "just patch it" would have felt obviously cheaper.
Optimize for blast radius when you're unsure. I couldn't prove Centrifugo would be faster to ship. I could prove a mistake in it couldn't break trades, alerts, or anything else on the platform.
Keep the escape hatch. The live switch back to the old server cost almost nothing to build and turned a one-way door into a two-way one. That made the decision much easier to get approved, and much easier to sleep on.
I used the same pattern earlier on BRDG, where every real-time message also has a polling fallback. Real-time is a performance feature; correctness should never depend on it.