← All writing

Written October 2026, looking back at Research4 min read

The missing 300ms

My real-time multimodal model for scoring a social feed was lagging by a few hundred milliseconds, and measuring every stage across 72,830 calls showed the cause wasn't where I expected.

From Solana and pump.fun · 2025 – 2026

After the execution engine, I built a research pipeline around the other half of the problem: a real-time multimodal model that scores a social feed as a market signal. A post comes in, possibly with images, the model decides whether it matters, and if it does, something downstream acts on it. Speed matters a lot here, because the value of a signal decays fast.

It was lagging by around 200 to 300 milliseconds behind where I thought it should be. This post is about finding out why, and why my first guess was wrong.

The model

Some context first. At the time, the model was a fine-tuned 7B vision-language model, trained with LoRA, quantized to 4-bit, and served with llama.cpp on AMD GPUs. (It later became a smaller fine-tuned 2B, which I describe in nights and weekends on Solana.) It returns a fixed JSON shape. I had a grammar to enforce it, but the fine-tuned model produced valid JSON on its own, and the grammar cost real time per token, so it stayed off on the hot path.

That setup replaced an earlier chain of four calls to a 70B model. One call to a small, specialised model doing the whole job was faster, simpler to operate, and, once fine-tuned, good enough. The four-call chain was the first big latency win. I assumed what was left was model inference or submit time.

Measuring every stage

Instead of guessing, I instrumented every stage and looked at 72,830 model calls.

My first suspect was the submit path, since that's where network round-trips live. It wasn't it. Metadata loading had a p50 of 10ms, and transaction build had a p50 of 17ms. Not free, but nowhere near 300ms.

Then I split posts by type. Text-only posts were fine: p90 of 18ms waiting in queue. Posts with media were not. Their queue wait was a p90 of 534ms and a p99 of 7.8 seconds.

So the time wasn't being spent running the model. It was being spent waiting to run it.

Time spent waiting in the queue, by post type (ms)Across 72,830 model calls. Text posts barely waited. Posts with media waited a long time, and the tail was enormous.
  • p90
  • p99
Show data
kindp90p99
Text posts18ms–
Posts with media534ms7.8s

Queueing explains the tail

Once I grouped by queue depth at arrival, the picture became obvious. Posts that arrived to an empty queue waited a p90 of 82ms. Posts that arrived with four items ahead of them waited a p90 of 2.2 seconds.

Queue depth at arrival vs p90 wait (ms)Arriving to an empty queue versus arriving behind four other posts.
Show data
depthp90 wait
Empty queue82ms
Four ahead2.2s

That's textbook queueing. Wait at arrival is roughly the number of items ahead times their service time, so four media posts at about 500ms each is about 2 seconds. The nonlinear part is utilization: as arrivals approach service capacity, the queue itself grows sharply, and long, variable service times like media make it worse. A burst of image-heavy posts would back everything up, and the text posts stuck behind them paid for it too. The averages looked tolerable. The tail was where the signal went to die.

What made media slow was a vision preprocessing pass running on the CPU, at roughly 420ms per post. The GPU was sitting there being fast, and the CPU was the bottleneck in front of it.

The fix

Two changes:

  • move that ~420ms vision pass from the CPU to the GPU
  • add separate lanes for media posts, so a burst of images can't block text posts behind it

The lanes are the change I'd underline. Speeding up the slow work helps, but isolating it stops the slow class of work from setting everyone's latency.

Before and after: separate lanesA burst of image posts used to back up everything behind it. Now media has its own lanes and the vision pass runs on the GPU.

Smaller things I learned along the way

One feature extractor. Training and serving share the exact same feature extraction code. Two copies that drift slightly apart is a quiet way to make a model look great offline and worse live, and I didn't want to find that out in production.

Libraries lie sometimes. The library I used to check whether the GPU supported bf16 said yes on one AMD GPU generation where it really didn't behave. Now key settings are chosen from the GPU architecture directly, not from a capability check I can't fully trust.

Most signals didn't survive. I replayed 201,796 posts against what actually happened next. Three of the four signals that looked promising didn't hold up, so I didn't change any thresholds based on them. That's a slightly deflating result, and also exactly why you replay before you tune. I wrote more about that instinct in running my own book.

What I took from it

Measure the stage, not the system. "It's slow" was true and useless. "Media posts wait in a queue, and the wait depends on depth" pointed straight at the fix.

Look at the tail, not the average. The medians looked fine the whole time. The p99 was 7.8 seconds.

And trust queueing theory over intuition. My intuition said the model was slow. The model was the fastest thing in the pipeline. What was slow was everything waiting for it.

Queue depth at arrival is now one of the first numbers I record in any pipeline. It's one number per request, and it explains the tail better than almost anything else.