← All writing

Systems3 min read

Designing for the retry

Any code that moves money will eventually run twice. The design rules I use so that "twice" is always harmless, with the state machines and diagrams behind them.

From Print.World · Aug 2025 – Present

Every system that moves money has retries in it: a client retries a request, a job runner retries a task, a network library resends a packet. Retries are good. They're also the most common way a correct system does something twice.

So when I design a money path, I start with one question: what happens if this runs again after the money already moved? This post is how I answer it.

The shape of the problem

The dangerous pattern isn't a crash. It's a sequence where every individual piece is reasonable:

How a correct system sells twiceThe side effect succeeds. A slow, non-critical step after it fails, and a retry wrapped around the whole operation runs the side effect again.

Retrying a failed order is reasonable. Looking up the trade afterwards is reasonable. A timeout on a slow query is reasonable. The problem is the composition: a retry boundary drawn around work that already partly succeeded, and an error from a non-critical step allowed to cross it. It's a well-known shape in trading systems, and it's the class of bug I design against from the first line.

Five design rules

1. Draw the retry boundary at the side effect. Retry the thing that hasn't happened yet. Once money has moved, everything after it is bookkeeping, and bookkeeping failures get logged and reconciled, never retried by re-running the trade.

2. "I don't know" is a state. On the perps work, a close or a take-profit/stop-loss order that fails without a clear answer from the venue isn't treated as failed. It's stored as UNKNOWN, and a retry replays the stored outcome instead of acting again. An authoritative rejection from the venue itself, on the other hand, means the order was never accepted, so it's safe to try again. Reconciliation works because every order carries a client order ID, so the system can ask the venue "did you get order X?" instead of guessing.

venue says ok        -> PLACED    (done)
venue says refused   -> REJECTED  (nothing sent, safe to retry)
no clear answer      -> UNKNOWN   (reconcile, never resend blindly)

3. Some things expire, and must stay expired. For cross-chain deposits to a single-use address, I made sure the system never re-broadcasts a deposit after its deadline. Re-sending late opens a window where two deposits can land.

4. Resending the same bytes is safe. Rebuilding is not. On Solana, rather than rely on SDK helpers that resend while they poll (which can report a trade that already landed as a failure), I send once, poll confirmation myself, and allow at most one guarded resend of the identical signed transaction. Identical bytes carry an identical signature, so they can only land once. Building a fresh transaction to "retry" is how you buy twice. I wrote more about this in nights and weekends on Solana.

5. Status is a version. On BRDG, a transfer row's status acts as its version number, and a writer only updates the row if the status is still what it read. Two pollers racing on the same transfer can't overwrite each other: a slow poll can never stamp a received amount onto a row that has already moved to REFUNDED. This works because those transitions don't revisit a status; where one can (an UNKNOWN that goes back to pending), a separate version number that only increases is the safe design.

Where the retry boundary belongsRetries live before the side effect. After it, failures are recorded and reconciled.

What I took from it

The scariest bugs in trading systems aren't crashes. Crashes are loud. These are quiet: the system does exactly what each piece was told to do, and the sum of it is a second sell nobody asked for.

So I don't ask "what happens if this fails?" I ask "what happens if this fails after the money moved?" That one question shapes more of my code than any library.