← All writing

Written October 2026, looking back at Systems9 min read

Sixteen months at IBM

CI/CD for Watson, then telemetry, then building Watson Code Review, an LLM review service on RabbitMQ and OpenShift. KV caching, when to re-run a review, how results reach developers, and how it scaled.

From IBM · May 2024 – Aug 2025

My IBM internship ran from May 2024 to August 2025, remote out of Markham, and it went in three stages: CI/CD for Watson, then telemetry, then building Watson Code Review, a service that reviews pull requests with an LLM. Each stage handed me a bigger piece of the system than the last.

  1. Stage 1CI/CD for WatsonBuild, test, and release pipelines for Watson services on OpenShift, plus an SDK and CLI for spinning up test environments.
  2. Stage 2TelemetryOpenTelemetry across services, error-biased sampling, and performance work on the agent that collects it all.
  3. Stage 3Watson Code ReviewAn LLM reviewing pull requests, from a side project to a service other teams relied on.

Stage 1: CI/CD for Watson

I started on the pipelines that build, test, and ship Watson services. A lot of engineers depend on those pipelines working, and almost nobody thinks about them until they don't.

The work was about making releases boring:

  • One pipeline system. I moved repositories from Travis CI onto Jenkins, carrying build history and artifacts across, without a window where teams couldn't ship.
  • Disposable test clusters. Integration tests ran on lightweight Kubernetes (K3s) clusters created per run, with the dependency images they needed managed by the pipeline instead of by hand.
  • Faster feedback. Parallel stages and caching so CI stopped being the slowest part of everyone's day.
  • Environments on demand. Engineers needed test VMs and OpenShift clusters from IBM's internal provisioning cloud, mostly through copied scripts. I owned an inner-source Go SDK for that API and a CLI on top of it, so an environment became one command.
A change on its way to releaseEvery change builds once, runs against a disposable cluster, and is promoted through environments by the same pipeline.

The lesson: infrastructure work is trust work. A pipeline that works is invisible, and that's the goal.

Stage 2: Building out telemetry

Once I knew how the services shipped, I moved to understanding how they behaved. I added OpenTelemetry tracing across Watson services so a single request could be followed through every hop it made.

The interesting decision was which traces to keep. Uniform sampling throws away the rare failures you actually need, so I pushed for error-biased tail sampling: keep every failing or slow trace, and sample healthy ones down. I wrote about that in detail in keep the failing traces.

I also worked on the agent that collects and exports telemetry. Many producer threads handing records to one exporter is a classic contention point, so the handoff became a lock-free queue, record layout was made cache-friendly after profiling with perf and VTune, and collectors got failover so a crashed primary hands over to a standby instead of dropping data.

The telemetry agentProducers never wait on a lock, and a standby collector takes over if the primary dies.

That telemetry ended up being the foundation for the next stage: I couldn't have run an LLM service responsibly without it.

Stage 3: Watson Code Review

Watson Code Review started as something I built on the side: a GitHub integration where a fine-tuned Granite code model, served on watsonx, reviews a pull request and leaves comments on it. It was later adopted as its own product inside IBM.

The model was the easy part to describe. The engineering was everything around it: getting large diffs through a fixed context window, keeping GPU time affordable, deciding when to review again, getting results back to developers in a form they'd trust, and scaling the whole thing on OpenShift.

The architecture

Watson Code Review, end to endEach stage is its own set of workers behind a RabbitMQ queue, so every stage scales and fails independently.

The service was split into stages connected by RabbitMQ queues, each with its own pool of workers (handlers):

  • Ingest receives GitHub's webhook, verifies its signature, ignores deliveries it has already seen, records a review job in PostgreSQL, and returns immediately. GitHub expects a fast response, and nothing slow belongs on that path.
  • Diff prep fetches the changes, splits them into chunks, and builds the prompts.
  • Inference sends chunks to the model, batching requests together.
  • Publish turns model output into comments and status updates on the pull request.

Splitting it this way meant each stage could scale on its own and fail on its own. A burst of large pull requests backs up the inference queue, not the webhook endpoint. A GitHub API hiccup backs up publishing, not the GPUs.

A few rules every handler followed:

  • Acknowledge after the work, not before. A message is acked only once its result is safely written, so a worker crashing mid-task means the message is redelivered, not lost.
  • Idempotent by design. Work is keyed on repository, pull request, and commit SHA, so a redelivered message finds the result already there and does nothing.
  • Bounded retries, then a dead-letter queue. Transient failures retry with backoff; anything that keeps failing goes to a dead-letter queue for a human, instead of looping forever.
  • Limited prefetch. Each worker only holds a few unacknowledged messages, so work spreads across workers instead of piling onto whichever one connected first.

Fitting a diff into a context window

A model has a fixed context window. Send a large diff whole and the tail is silently cut off, which means part of the change never gets reviewed and nobody knows.

So diff prep splits every pull request by file and by hunk into chunks that fit a per-model token budget, keeping enough surrounding lines for the model to understand each change. A large pull request becomes several requests instead of one truncated one.

KV caching: making the shared part free

This is the part I found most interesting.

Every review prompt had a large part that was the same across requests (the instructions, the review guidelines, the output format) and a small part that was unique (the diff chunk). If the unique part came first, nothing could be reused. So the prompt layout was designed around the cache:

Prompt layout, designed for prefix cachingThe most shared content goes first and is kept byte-for-byte identical, so its prefill is computed once and reused. Only the tail differs per request.

A few details made the difference:

  • Most stable first. Global instructions, then per-repository guidelines, then file context, then the chunk. Each layer is shared by a smaller group of requests, so the cached prefix is as long as possible for each.
  • Byte-for-byte identical. Anything that varies, like a timestamp or a reordered list, breaks the prefix. The shared parts were generated deterministically so the cache actually hit.
  • Send related chunks to the same worker. Chunks from the same pull request share the most prefix, so routing them together means the cache on that worker is already warm.

Combined with dynamic batching, where a worker collects requests for a few milliseconds and runs them through the GPU together, this is what made the service affordable. Batching trades a small amount of latency for much better GPU throughput, and the wait is capped so no single review sits waiting for a batch to fill.

When to review again

Pull requests change. Deciding when to re-run a review mattered as much for cost as for usefulness:

Deciding whether to re-runA new push supersedes any review still in progress, and only chunks that actually changed are reviewed again.
  • Supersede, don't stack. A new push makes any in-progress review of an older commit obsolete, so it's cancelled rather than finishing and posting stale comments.
  • Debounce bursts. Developers often push several commits in a row; waiting a moment avoids reviewing each one.
  • Review the change, not the whole PR. Chunks are identified by their content, so a chunk that hasn't changed since the last review reuses its earlier result.
  • Skip what shouldn't be reviewed. Draft pull requests, generated files, and lockfiles are skipped, and a developer can ask for a fresh review explicitly when they want one.

Getting results to developers

A review is only useful if developers trust it, and trust comes from clarity:

  • A status on the pull request that moves from queued, to in progress, to complete, so nobody wonders whether it's running.
  • Inline comments anchored to the exact lines they're about, plus one summary comment.
  • Update, don't duplicate. On a re-run, earlier comments are updated or resolved rather than reposted, so a pull request doesn't fill up with repeats.
  • Honest failures. If a review couldn't run, the status says so plainly, instead of silently showing nothing.
  • Traceable. Every review carries its trace ID, so when someone asks "why did it say that?", I could follow that exact request through every stage.

Scaling it on OpenShift

Each stage ran as its own deployment on OpenShift, so it could be scaled and rolled out independently.

Scaling on queue depth, not CPUQueue depth is the honest measure of work waiting, so it drives scaling. GPU workers scale within a fixed budget, and backpressure keeps the queue bounded.
  • Scale on backlog. CPU usage is a poor signal for a queue-driven service: a worker waiting on the model looks idle while work piles up. Queue depth is the honest signal, so it drove scaling.
  • Different limits per stage. The CPU-bound stages scale wide cheaply. GPU inference scales within a fixed budget, so when the budget is reached, work queues instead of costs growing without limit.
  • Lanes by size. Small pull requests shouldn't wait behind a huge one, so work was separated by size, the same idea I later used for text and media lanes in my own ML pipeline.
  • Graceful shutdown. On scale-down or deploy, a worker stops taking new messages, finishes or returns what it holds, and only then exits, so rollouts don't lose work.
  • Health checks that mean something. A worker reports ready only once it can reach the queue and its downstream dependencies, so traffic never goes to a pod that can't do the work.

What working at that scale felt like

The first few months were humbling. On my own projects I owned everything. At IBM I owned a small piece of something many people depended on, and most decisions involved people I'd never met.

  • Write things down. On a remote team across time zones, a decision that isn't written down didn't happen.
  • Reviews are a conversation. I started treating code review as the cheapest way to borrow someone else's experience.
  • Scope is negotiated. "What's the smallest version of this that helps someone this month?" saved me more time than any tool.
  • Ownership includes the boring parts. Owning a pipeline, an SDK, or a service means answering the message about why it's broken, whenever it arrives.

What's next

While I was at IBM, my evenings went somewhere else entirely: Solana. And as this internship wound down, I started at Print.World. Both made more sense because of what IBM taught me about owning something other people depend on.