Written October 2026, looking back at Markets5 min read
Racing relays without paying twice
An independent execution engine that turned social signals into on-chain trades, and how it raced several relays at once while guaranteeing only one trade lands.
From Solana and pump.fun · 2025 – 2026
Outside of work I built an execution engine that turned social signals into on-chain trades on Solana, as part of my pump.fun launch project. It was independent of Print, and it ran against live mainnet traffic, so the design was tested by reality rather than by me. The signal side is its own story (see the missing 300ms). This post is about the execution side: getting a trade on-chain in milliseconds, without accidentally doing it twice.
Racing relays
To get a transaction included quickly, you submit it to block builders through relays. No single relay wins every time, so the obvious move is to send to several in parallel and let the fastest one land.
The catch is that each relay wants its own tip, paid to its own account, inside the transaction. So you can't send one transaction to everyone. You build one transaction per relay, each with a different tip instruction.
Which raises the real problem: if you send three different transactions that each buy the same thing, what stops two of them landing?
One winner by construction
For repeat trades, every copy uses the same nonce account instead of a recent blockhash. A nonce-based transaction is valid only while the nonce holds a specific value, and landing it advances the nonce. So whichever copy lands first advances the nonce, and every other copy instantly becomes invalid.
There's no coordination, no cancel message, no race to clean up. Exactly one copy can succeed because of how the transactions are built.
The trade-off to design around: a nonce transaction never expires on its own. If no copy lands, every copy stays valid until the nonce moves, so a stale copy could still land later at a worse price. Two things protect against that: every buy carries a hard ceiling on what it's allowed to spend, and advancing the nonce yourself cancels every outstanding copy at once. The advance-nonce instruction also has to be the first instruction in the transaction.
The next issue was rapid back-to-back trades. If two different trades share one nonce account, the first one invalidates the second. So there's a small pool of nonce accounts, handed out round-robin, so consecutive trades don't collide while copies of the same trade still do.
On top of that, each relay has several regional endpoints, and I fire at all of them at once. Every submission logs which endpoint and region won and how long it took, which is how I learned where to deploy.
Taking round-trips out of the click path
The second half of speed is not waiting for things you could have done earlier. The rule I used: if a step can happen before the signal arrives, it should.
- a pool of pre-fetched blockhashes, so grabbing one takes under a millisecond when warm
- keep-alive pings so connections are already open
- metadata, blockhash, and lookup-table loads run in parallel instead of in sequence
- return as soon as the transaction is submitted, and confirm in the background
- deploy the engine physically near the relays
Warm state, shared on purpose
The first version ran as Next.js API routes, which is not where a latency-critical engine belongs, and it showed: Next.js bundles each route with its own copy of every module it imports. That's fine for stateless code and quietly wrong for warm pools: a pool warmed in one route is invisible to another.
So the warm state (blockhash pool, relay connections, nonce pool) lives on globalThis, as a single shared instance every route reads from, and routes only pass plain data between each other. Every submission also logs which relays it actually raced and which one won, so "are we racing three relays or one" is a number on a dashboard, not an assumption. Today I'd run the engine as one long-lived process and skip the problem entirely.
Fitting in 1,232 bytes
Solana transactions have a hard size limit of 1,232 bytes. A swap plus tip plus nonce instructions plus all the accounts gets close fast. Each account address is 32 bytes. With a lookup table, the transaction can reference an account by a 1-byte index instead, which made the difference between fitting and not fitting.
- Inline addresses
- Via lookup table
Show data
| accounts | Inline addresses | Via lookup table |
|---|---|---|
| 10 accounts | 320 | 10.00 |
| 20 accounts | 640 | 20.00 |
| 30 accounts | 960 | 30.00 |
Compute limits follow the same philosophy. The priority fee is charged per unit of the compute limit you request, so an oversized limit is money for nothing, and an undersized one fails on-chain. Limits come from measured usage plus a fixed headroom, not from a guess at what the code should cost.
Safety
Because this moved real money, demo and real executors are completely separate code paths, and every spending action checks at runtime that the demo path can never send a real transaction. I didn't want the difference between a test and a real trade to be one flag someone forgets.
The EVM side
I also supported EVM chains, where the problems look different. Nonces there are per wallet and strictly sequential, so I kept a local nonce manager per wallet behind a lock, and resynced from chain whenever a node said "nonce too low". Submission went to several RPCs at once, first success wins, which is the same idea as racing relays with simpler plumbing.
What I took from it
Two things. First, build the "only one can win" guarantee into the data, not into coordination logic. The nonce trick is robust because there's nothing to get wrong at runtime. Second, make the race observable. Logging which paths every trade raced, and which one won, is what turns "it feels fast" into a design you can reason about and tune.