Ramjet vs NVIDIA Dynamo: Which Router for Coding-Agent Traffic?
Sep 30, 2026
On one 8×H200 server running GLM-5.3, Ramjet kept 92% of agent prompts in cache against 83-85% for NVIDIA Dynamo 1.5.0's cache-aware router, and completed 22-29% more agent turns a minute on the same engine. Dynamo is built for larger fleets, where we expect it to pull ahead.
Yesterday we wrote about serving the full GLM-5.3 to 48 coding agents from one 8×H200 server. Most of that gain came from Ramjet, our open-source load balancer, sending each agent's next turn back to the GPU that still held its conversation. NVIDIA Dynamo does the same job differently. Its router does not have to guess where a conversation is cached, because the engines report it.
We ran both in front of the same engine on the same server. With identical engine software and settings, Ramjet completed 22% more agent turns a minute at 32 developers and 29% more at 48, with the same wait for the first token. The difference follows the cache. Ramjet served 92% of prompt tokens from memory; Dynamo served 83-85%.
One server, one run per cell. The third row runs Ramjet on the exact engine build Dynamo ships.
That is one server. Dynamo is built for fleets: several routers, servers that come and go, and a cache shared between machines. We did not test those, and past a handful of servers they are likely to decide the question in Dynamo's favour.
Two ways to find a conversation's cache
A coding agent re-sends its whole conversation on every turn. GLM-5.3 on this server runs as eight independent ranks, one per GPU, each with its own cache of recently read prompts. If a turn reaches the rank that read the previous one, the engine only processes the new tokens. If it lands elsewhere, the engine reads tens of thousands of tokens again.
Ramjet guesses where each conversation lives. It fingerprints the start of every request in 2 KB chunks, remembers which rank served each chunk, and weighs that overlap against how busy each rank is. It never talks to the engine, so it cannot see a conversation being evicted. It runs as one binary with no other services.
Dynamo's router is told. Each engine rank publishes an event whenever it stores or drops a block of cached prompt, including blocks moved to host memory. The router keeps an index of every block on every rank. It scores each rank by the prompt work left after crediting what is already cached, plus the answer-generation load already running there. A block in host memory counts at 0.75 of a block on the GPU.
Dynamo's front end also handles chat templates, tokenization and tool-call parsing in Rust, work that otherwise runs in the engine's Python process. Across machines it needs etcd for discovery, plus NATS if you run more than one router.
The test
The engine configuration is the one from yesterday's post. SGLang splits attention across eight ranks with an 8-bit cache. DeepEP handles the expert exchange, multi-token prediction speeds up generation, and each rank has 32 GB of host memory as a second cache tier.
Dynamo 1.5.0 ships its own build of SGLang, version 0.5.18. Our deployment runs 0.5.20 with its hrrn scheduler, which picks the next waiting request by how long it has waited relative to how much work it needs, so short requests do not queue behind long ones. To keep the router and the engine version apart, we ran four configurations on the same server:
| Configuration | Router | Engine |
|---|---|---|
| Dynamo | Dynamo 1.5.0, --router-mode kv, per-rank cache events | SGLang 0.5.18 from ai-dynamo/sglang-runtime:1.5.0 |
| Dynamo, tuned | the same with --router-kv-overlap-score-credit 2.0 | the same |
| Ramjet, same engine | Ramjet v0.7.0, one upstream per rank | the same SGLang 0.5.18, launched directly |
| Ramjet, current | Ramjet v0.7.0 | SGLang 0.5.20 with hrrn |
The workload is the same simulated team as yesterday (agent_swarm_bench.py). Each developer drives a coding agent that is almost never idle, and prompts average 35,000-43,000 tokens. Each cell ran for 10 minutes after a 2-minute warm-up, with fresh prompts so no configuration inherited another's cache.
Results
| Configuration | Developers | Turns / min | Cached | ms per generated token, median | Wait for first token, median / p90 |
|---|---|---|---|---|---|
| Dynamo | 32 | 73.0 | 85.0% | 71 | 2.06 s / 7.7 s |
| Dynamo, tuned | 32 | 69.7 | 83.4% | 78 | 2.36 s / 5.9 s |
| Ramjet, same engine | 32 | 88.9 | 92.4% | 52 | 1.53 s / 8.7 s |
| Ramjet, current | 32 | 95.9 | 92.4% | 48 | 1.28 s / 7.5 s |
| Dynamo | 48 | 77.3 | 84.7% | 107 | 2.92 s / 20.0 s |
| Dynamo, tuned | 48 | 82.5 | 84.9% | 102 | 3.01 s / 19.2 s |
| Ramjet, same engine | 48 | 99.5 | 91.7% | 73 | 2.88 s / 20.5 s |
| Ramjet, current | 48 | 102.8 | 92.0% | 66 | 2.99 s / 20.3 s |
No request failed in any configuration.
The router accounts for most of the gap. On the identical SGLang 0.5.18 build, Ramjet completed 22% more turns at 32 developers and 29% more at 48. Moving Ramjet to SGLang 0.5.20 added a further 8% and 3%.
Each point of cache Dynamo missed meant re-reading part of a long conversation. That extra reading also slowed answer generation for every other agent on the same GPU. Generating a token took 71-107 ms under Dynamo against 52-73 ms under Ramjet on the same engine.
Dynamo's tuned setting gave the shortest tail of any configuration at 32 developers: the slowest one in ten turns started within 5.9 seconds, against 7.5-8.7 seconds with Ramjet. It spreads queues evenly across ranks, at the cost of cache hits.
What we ruled out, and what is still open
The gap is not Dynamo ignoring the ranks. It spread requests evenly, 87 to 131 per rank in the first cell. It is not blind to host memory: 19% of its routing decisions counted blocks held there. Its block size matched the engine's 64-token pages. Doubling the credit it gives cached blocks did not raise the hit rate.
We have two unverified explanations. GLM-5.3's sparse attention keeps a second, smaller index cache beside the main one, and the router's view of the two may not line up. Or Dynamo's load weighting moves too many sessions when nearly every request carries a large cached prefix. Someone who knows Dynamo's scoring would find the cause faster than we can.
Limits of this comparison
- One run per cell. The configurations shared a server but ran one after another, Dynamo first. Between two servers of the same type we have seen up to 10% difference at high load.
- One model and one workload. Agent traffic re-sends long, highly reusable prompts, which is the case prefix routing is designed for. Traffic with short, unrelated prompts depends more on load balancing, where Dynamo's exact view of each rank should do better.
- One engine version. A Dynamo release built on a newer SGLang could change the picture.
- One tuning knob. Dynamo can also pin sessions by a session ID, randomise placement, and change its queueing policy. We tried none of these.
- One server. Eight ranks behind one router is Ramjet's strongest case.
- We wrote Ramjet. The harness, deployment files and raw numbers are public so anyone can rerun them.
Where Dynamo should win: more than a handful of servers
Nothing above measures a fleet. Somewhere past about five 8-GPU servers, 40 or more ranks behind one router, we expect Dynamo's design to pay for its extra parts. That threshold is our judgement, not a measurement.
Guessing gets harder with more replicas. Ramjet's fingerprints work because one process sees every request. We simulated its router on synthetic sessions across growing fleets. Its current scoring keeps 90% of turns on the right replica at 40 replicas. The older scoring fell to 51%. A router that reads cache events from the engines does not have to guess, so it should not degrade the same way.
A simulation of Ramjet's routing code on synthetic sessions, with no engines and no measured cache.
One Ramjet is active at a time. Ramjet keeps its routing state in memory, so a second copy can only stand by, and it starts from an empty map after a failover. Dynamo runs several routers at once. Each rebuilds the cache index from the same event stream, and --router-replica-sync shares in-flight load between them over NATS.
Servers come and go. Adding a server to Ramjet means editing a topology file and a four-second restart. Dynamo workers register themselves in etcd or Kubernetes and appear in the router while it runs. That matters with autoscaling, which Dynamo's Planner provides, or with capacity that can be reclaimed.
A cache shared between servers. Dynamo understands a Mooncake store behind SGLang's host-memory tier and credits blocks any worker can fetch over the network. A conversation that has to move, after a failure or to relieve a busy server, then pulls its cache from another machine without re-reading it. Ramjet can only try to avoid moving it.
A separate pilot of ours points the same way. On two servers, routing each session to the right server but letting the engine choose the rank cut cache hits to 65%. On a single server, a Mooncake tier with half the host memory beat the full host tier on its own: 11% more turns and 94.7% cached against 91.9%.
Splitting prompt reading from answer generation. With RDMA networking between servers, Dynamo can run prompt reading and answer generation on separate machines and move the cache between them. Agent traffic has little fresh prompt to read, so this helps less than it sounds. It would help a low-latency tier: on our server, generation slowed from about 43 tokens a second per request at 16 developers to 13-16 at 48-64, partly because prompt reading interrupts it.
Tokenization scales with the router. One SGLang tokenizer process saturated on 100,000-token agent prompts, so we run four. Dynamo tokenizes in its front end, so adding routers adds tokenizer capacity.
The cost of all this is operational. Dynamo's documentation recommends a three-node etcd cluster for production, NATS if routers are replicated, and Dynamo's own engine build, often with its Kubernetes operator. Ramjet is one binary in front of an unmodified engine.
Which to use
| Situation | Our choice |
|---|---|
| One to five servers, agent or chat traffic with a lot of repeated prompt, an unmodified engine and few moving parts | Ramjet. 22-29% more turns on one server in our test |
| You care most about the slowest turns on mixed traffic | Test both. Dynamo had the shortest p90 wait in our tuned run |
| More than about five servers, or several active routers | Dynamo. Exact cache view across the fleet, replicated routers, live membership |
| Autoscaling or capacity that can be reclaimed | Dynamo, with its Planner and live discovery |
| An RDMA network and a shared cache or split prompt and answer servers | Dynamo with Mooncake |
| Short prompts with little reuse | Either. The cache advantage disappears |
| Many self-contained servers on plain Ethernet | A stateless gateway that picks the server, with Ramjet choosing the rank inside it, or Dynamo as one system |
We will keep Ramjet on each server, where it measured faster. Before we grow past a few servers we will rerun this comparison on two or three of them, with Mooncake on, repeated cells and the order alternated.
Reproducing it
- Engine and Ramjet:
deploy/glm53_h200in the Ramjet repository. - Workload:
bench/agent_swarm_bench.py BASE_URL glm-5.3 --developers 32 --duration 600 --warmup 120. - Dynamo:
nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.5.0, run as root because the image's default user cannot write the compile cache. Front end:python3 -m dynamo.frontend --router-mode kv --discovery-backend file. Worker:python3 -m dynamo.sglangwith the same engine flags plus--dyn-tool-call-parser glm47 --dyn-reasoning-parser glm45and a ZMQ--kv-events-config. - Every number above is in the Ramjet experiment journal under 30 September.
Ramjet is open source under Apache-2.0. The engine setup behind these numbers is in Serving Full GLM-5.3 to 48 Coding Agents From One 8×H200 Server. Talk to Helix about running GLM-5.3 on your own hardware.