← All writing

Markets12 min read

Nights and weekends on Solana

How a side interest during my IBM internship turned into an automated pump.fun launch pipeline, fed by a paid real-time X feed, from a post on X to a landed Solana transaction, measured stage by stage.

From Solana and pump.fun · 2025 – 2026

During my IBM internship, my days were Go, CI/CD, and Watson. My nights, starting around spring 2025, were Solana. This is the story of how that side interest grew into the most demanding system I've built on my own: an automated pump.fun launch pipeline that reads posts from X as they're published, decides with a local model whether one is worth acting on, and lands a transaction on-chain fast enough for that to matter.

If you've never touched crypto, the primers along the way cover what you need.

Learning the chain

I started the way most people do: small programs and a lot of reading. In spring 2025 I wrote a few toy projects, then worked through Blueshift's challenges in May (an Anchor vault and escrow, then the same vault in Pinocchio, which strips away Anchor's conveniences so you see every account check), and built a Rust bot that provided liquidity on fresh pools, collected fees, and left before its position went stale.

The moment it clicked was contributing upstream. In June 2025 I got two changes merged into Typhoon, a Rust framework for Solana programs. The first added inline hints to account validation, which Criterion benchmarks showed cut a single validation from about 3.59ns to 3.26ns. The second inlined trait methods and macro-generated code, saving over a hundred compute units across instructions. Small diffs, but they meant reading someone else's codebase until I understood why every compute unit was spent where it was.

2Merged upstream changes to Typhoon
128+Compute units saved across instructions
  1. May 2025Blueshift challengesAn Anchor vault and escrow, then the same vault in Pinocchio with every account check written by hand.
  2. June 2025Upstream to TyphoonTwo merged changes: inlined account validation and tighter macro-generated code.
  3. July 2025Turbin3 builder cohortVault, escrow, an AMM, an NFT marketplace and staking.
  4. Aug 2025A yield vault programDeposits rebalanced across lending protocols through protocol adapters.

Solana rewards that kind of attention. Every transaction has a hard size limit, a compute budget, and a short window before its blockhash expires. You can't paper over a slow path with more servers. That's what pulled me in.

The first sender

My first real Solana system, in May and June 2025, was a transaction sender: build a transaction, sign it, and get it into a block as fast as possible. That sounds simple until you learn how blocks actually get built.

How a Solana transaction landsA signed transaction has to reach whichever validator is the leader right now, before its recent blockhash expires, and win a place in that leader's block.

A transaction is sent to the current leader. Under load, though, the leader can't accept everything, so it prioritizes. Two mechanisms matter:

  • Fees and tips. You can attach a priority fee, and relays accept a separate tip, to move up the queue.
  • Stake-weighted quality of service (SWQoS). A leader reserves most of its incoming connection capacity for connections from staked validators, shared out in proportion to their stake. RPC nodes and relays get that priority by peering with staked validators. If you send from an unstaked connection during a busy slot, your packets are the first to be dropped.

So fast senders don't go direct. They go through relays that hold staked connections or feed block builders, and pay a small tip for priority. No single relay wins every slot, so you send to several at once. That raises the problem I wrote about in racing relays without paying twice: if you send the same buy three ways, how do you guarantee only one lands? (Short answer: a durable nonce, so the first to land invalidates the rest by construction.)

The pump.fun trade

From 2025 into 2026, my main project was built around one observation: on Solana, attention had become a market.

That design turned engagement into something you could trade. When a post goes viral on X, tokens named after it appear within minutes, and people trade them on the story. The first coherent token, with the right name, ticker, and image, tends to capture most of the attention and most of the flow.

To me that looked like a market ripe for automation, because every step was mechanical:

  • The input is public. Posts on X are visible to anyone, the moment they're published.
  • The decision is classification. Is this post launchable? What's the obvious name, ticker, and image?
  • The payoff is speed. Being first with a coherent launch matters more than anything else.

So the whole game came down to two questions: how fast can I see a post, and how fast can I turn it into a landed launch?

Paying for speed: a real-time X feed

Most people see a post when their app refreshes, seconds after it was published. X's standard API wasn't built to push posts from specific accounts to you the instant they appear. So I paid for websocket access to a real-time X data feed: I give it the list of accounts I care about, and it pushes their posts to me over a persistent connection as soon as they're published.

Measured that way, text posts arrived on my server about 200ms after they were published, at the median. A few design choices kept it there:

  • Several connections, first arrival wins. The same feed ran over several connections in different regions at once. Whichever copy of a post arrived first was used; the rest were dropped by a deduplicator keyed on the post ID.
  • Stamp time on arrival, not after processing. The receive time is recorded the moment the bytes land, before parsing, so every later latency number starts from the same clock.
  • Don't wait for what you don't need. Text is enough to decide most launches. Images are fetched in parallel, under a strict time budget, so a slow image never holds up a text decision.
~200msFrom a post being published to it arriving on my server
268msMedian to build and submit the launch transaction
247msMedian model decision on a text post

That was the advantage: by the time most people had even seen a post, the pipeline had already read it, decided whether it was launchable, named it, and had a launch transaction on its way to a validator.

From a post to a launched tokenA paid real-time feed delivers posts from tracked accounts, models propose a name, ticker, and image, and the launch goes out automatically by rule.

I built the whole path: the feed consumer, the launch orchestrator, the transaction builder, the relay clients, and the blockhash and nonce pools.

Bundles: making a launch atomic

A launch isn't one transaction. On pump.fun the creator can buy in the same transaction that creates the token, and the opening buys follow right behind it. The engineering problem was making the create and those opening buys land together, so a launch could never end up half-built.

So a launch went out as a bundle: the create transaction first, then the opening buys right behind it, all in the same slot. A few design details made that work:

  • Pricing each leg in advance. Every buy is priced against where the bonding curve will be after the create and the buys ahead of it, with a ceiling on how much it's allowed to spend.
  • Tips and fees per the relay's rules. Block engines auction a bundle as a unit, mostly on its tip. The relay I sent bundles through also screened each leg against a fee floor and ranked a bundle by its lowest-paying leg, so every leg was raised to match the highest one. That's specific to that relay; through a block engine directly, one well-placed tip is the cheaper design.
  • Atomicity as a feature. Because a bundle is all-or-nothing, there's never a half-opened launch where the create landed and the buys didn't.
  • Verifying the result. A bundle either lands whole or not at all. There was also a non-bundle fallback that raced individual buys into the create's slot, and that path can land split across slots or partially, so the system checks which happened and acts on each case differently.
One launch, one bundleThe create and the opening buys execute in order in the same block, or not at all.

Fitting it into a transaction

  • Lookup tables. A Solana transaction is capped at 1,232 bytes, and every account address is 32 bytes. A lookup table lets a transaction reference an account by a 1-byte index instead, but only for accounts that already exist and don't sign: signers always cost a 32-byte key plus a 64-byte signature, and accounts derived from a mint created a second ago can't be in a pre-built table. Our own table carried the static accounts (programs, fee and config accounts) plus fixed recipient wallets, which freed enough room for the opening buys. Relays check for their tip account directly in the transaction, so tip accounts deliberately stay out of the table.
  • Compute budget. The priority fee is charged per unit of the compute limit you request, not what you use. So the limit is set from measured usage plus headroom, kept tight, and the fee follows from it.
  • Warm paths. A pool of fresh blockhashes, kept-alive relay connections, and metadata, blockhash, and lookup-table loads in parallel, so the click path does almost no network round-trips.

On the model side, the first version used hosted LLMs, switched at runtime through Redis, with prompts versioned and A/B tested. It worked, but every call went over the internet to someone else's GPU, and I kept wanting to know where the time went.

Bringing the models home

In mid 2026 I rebuilt the decision side around my own models, running locally, deciding which posts were launchable. This is where the machine learning came in properly.

The feed. Posts arrive over Redis pub/sub from a publisher service that holds the upstream connections. Each event is stamped with a monotonic receive time the moment its bytes arrive, before parsing or dedup, so every latency number starts from the same clock. Decoding the timestamp embedded in X's post IDs, text posts reach the pipeline around 218ms after they're published at the median.

The pipeline, end to endText and media take separate lanes so an image can never hold up a text post. Packets show the path a post takes.

The models. The first version replaced a chain of four calls to a 70B model with a single call to a fine-tuned 7B vision-language model. Later I fine-tuned and shipped a smaller 2B version, quantized and served with llama.cpp on two AMD GPUs on ROCm:

  • five llama.cpp backends, split into text lanes and media lanes, so an image can never block a text post
  • a small router in front that sends each request to the right pool by least-outstanding requests, benches unhealthy lanes, and only re-admits them after a passing probe
  • a priority queue that sheds anything older than a few seconds, because a stale signal is worth nothing
  • a LightGBM gate after the model, trained on tens of thousands of labelled posts, that decides whether a scored post is worth acting on, with its threshold calibrated on live traffic rather than set by hand
Where the time went, per stage (ms)Post-to-receipt is derived from tweet ID timestamps for text posts. Model and decide stages from 26,490 text and 20,109 media posts. Submit from 214 live sends.
  • p50
  • p90
Show data
stagep50p90
Post to receipt (text)218ms370ms
Model call (text)247ms393ms
Model call (media)537ms813ms
Decide (text)12ms17ms
Build and submit a transaction268ms580ms

The latency hunt that shaped this design is its own post: the missing 300ms. The short version is that the model wasn't slow; posts with images were waiting in a queue behind each other, and separate lanes fixed it.

Getting it on-chain

The execution side reused everything I'd learned, with a few new rules:

  • Several paths at once. Within one relay, the same signed bytes go to its global edge and a regional gateway in parallel, plus a plain RPC node, and the first acceptance wins. Across different relays, each needs its own tip, so each gets its own transaction, all built on one shared nonce so only one can land (that's the racing relays design).
  • Rebroadcast the same bytes, never a new transaction. If a send hasn't confirmed, the identical signed transaction is resent on a short schedule. Same signature, so it can land at most once. Building a fresh transaction to "retry" is how you buy twice.
  • Race two ways of confirming. Polling signature status over HTTP races a websocket subscription, and whichever answers first wins. The launch path waits at processed commitment for speed, accepting that a processed result can, rarely, be rolled back by a fork. An answer of "unknown" is never retried blindly, and rebroadcasting stops once the blockhash expires (about 150 blocks).
  • Watch the chain, not the API. A gRPC account stream sees tokens arrive in our own accounts, which is the real signal that a trade landed.
Sending one transaction several waysIdentical bytes to every endpoint of one relay. Because the signature is identical, at most one copy can ever land.

Over 214 live launches, the submit step took a median of 268ms (p90 580ms). Relay acceptance alone was about 35ms at the global edge, while the regional gateway ranged from about 46ms to over 500ms, and the slowest launches were the ones included a few slots late. What I never built was a clean per-relay table of land rate and slots-to-land, and that's the first thing I'd add: build-and-submit is a client metric, but landing in the right slot is the one that matters.

What it taught me

Solana taught me to treat latency as a budget you spend on purpose, stage by stage, instead of a number you look at afterwards. It also taught me that the scary bugs in fast systems aren't crashes. They're the transaction that lands twice, the pool that's silently empty, the model lane quietly blocked by an image. Every one of those I found by measuring each stage separately.

This was a personal project, separate from my work at Print.World. It's also why, when I joined Print.World in August 2025, the problems felt familiar: quotes, routing, and landing, with a bigger team and a lot more users behind them.