Ramjet on DeepSeek-V4.1-Flash: Routing Still Pays When the Cache Never Fills
Oct 7, 2026
On one 8×H200 server, DeepSeek-V4.1-Flash never filled more than 22% of its prompt cache. Sending each coding-agent session back to the engine that already read it still added 18% throughput. Here is what Ramjet did, and two ideas that did not work.
On our 8×H200 server, DeepSeek-V4.1-Flash runs as two copies of the model, each spread across four GPUs. Each copy keeps up to 4.9 million tokens of reusable prompt data in GPU memory (the key-value cache, or KV cache). Under our heaviest coding-agent load, the fullest copy used 22% of it. Nothing was ever evicted.
We expected that to make the load balancer in front of the two copies less important. On GLM-5.3, Ramjet earned its keep by keeping agent sessions next to a cache that was too small to hold every conversation. With a cache this large, any copy could hold everything.
It still mattered. Sending each agent session back to the copy that had already read its prompt was worth 18% more agent turns per minute at 64 developers. This post explains why, and which Ramjet setting made the difference. It then covers a second copy that serves the full one-million-token context, and a replica that looked healthy but answered nothing, which changed how Ramjet checks its engines.
The server and the workload
The model is deepseek-ai/DeepSeek-V4.1-Flash, served with SGLang 0.5.21 on one server with eight H200 GPUs. Each of the two copies spans four GPUs and answers on its own port. Its 196-billion-parameter Engram memory lives in host RAM, which is what lets a copy fit on four GPUs at all. Our earlier post on the architecture explains the encoder, Engram and the compressed cache.
The load comes from Ramjet's agent benchmark (agent_swarm_bench.py). Each simulated developer drives a coding agent that works almost continuously. Every turn re-sends the whole conversation: the harness's system prompt and tools, the repository instructions, and every message and tool result so far. Prompts averaged 37,000 to 44,000 tokens and reached about 120,000 before the agent compacted them.
So almost all of each prompt was already read on the previous turn. If the copy that receives the turn still holds that work in its cache, it reads only the few thousand new tokens. If the turn lands on the other copy, that copy reads the whole conversation from scratch.
With the setup described below, the server completed 131 to 140 agent turns a minute at 16 developers (across two runs), 217 at 32, 276 at 64 and 287 at 96. Between 90% and 94% of prompt tokens came from cache, and no request failed.
Routing back to a warm copy is worth 18%
To measure what routing is worth, we kept both engines running and changed only Ramjet's settings between five-minute runs. Each run used fresh prompt text, so no run could reuse another run's cache.
| Ramjet routing at 64 developers | Agent turns a minute | Prompt tokens from cache | Turns that stayed on their session's copy |
|---|---|---|---|
| Least loaded, ignoring the cache | 222.8 | 88.6% | 50% |
| Prefix affinity, absolute credit (default) | 222.7 | 88.5% | 60% |
| Prefix affinity, marginal credit | 263.8 | 91.6% | 88% |
| Prefix affinity, relative credit | 263.5 | 91.5% | 83% |
| Marginal credit with prefix single-flight | 263.6 | 92.0% | 87% |
The cache-hit rate moved by only three points, which understates the cost. Each turn that changes copy re-reads a conversation of about 40,000 tokens. With least-loaded routing the two copies read 15,700 fresh prompt tokens a second, against 14,000 with marginal credit. They did more prompt reading and still completed 15% fewer turns.
At 16 developers the gap was 7%: 129.6 turns a minute against 120.8, with the median agent turn finishing in 2.14 seconds instead of 2.33.
Why Ramjet's default setting lost
Ramjet scores each copy by how much of the incoming prompt it has already served, minus a penalty for how busy it is. The default, absolute credit, counts every matching block of the prompt's beginning, up to a cap.
Coding-agent prompts all start with the same large block: the harness's system prompt and tool definitions, often more than 64 KiB. Both copies have seen it thousands of times. Under absolute credit, both copies score the full cap for almost every request, so the score cannot tell them apart and the load penalty decides. Sessions hop between copies, and 60% stayed put.
Marginal credit counts only the blocks a copy holds beyond what the least-warm copy also holds. The shared system prompt cancels out, and the part that differs is the session's own conversation, which only one copy has seen. Relative credit compares each copy with the warmest one instead, which behaves the same with two copies and holds up better across a larger fleet. For agent traffic, set RJ_ROUTE_AFFINITY_BASIS=marginal with two replicas, or relative beyond that.
Ramjet's prefix single-flight, which sends concurrent first requests that share a new prefix to the same copy, made no measurable difference here. The benchmark rarely starts two sessions with the same new prefix at the same moment.
A second copy for one-million-token prompts
DeepSeek-V4.1-Flash supports a one-million-token context, but we serve our two copies at 262,144 tokens. The model's draft-and-verify decoding (DSpark) needs a large temporary buffer when the engine prepares for many simultaneous requests, and at the full context that buffer did not fit on our GPUs. The draft decoding roughly tripled single-request speed, so we kept it and shortened the context.
To offer the full context anyway, we reconfigured one copy for one million tokens, with draft decoding still on and fewer simultaneous requests prepared in advance. Ramjet's long-prompt lane then sends any request body over 1 MB (about 250,000 tokens) to that copy only (RJ_ROUTE_LONG_PROMPT_BYTES=1000000). Through Ramjet, the long copy recalled an eight-digit code planted a quarter of the way into a 463,313-token prompt in 74 seconds, and in an 834,736-token prompt in 207 seconds.
The long copy is slower at ordinary work. With one request it generated 312 tokens a second against 354 for a 262k copy. With 16 simultaneous requests it managed 75 tokens a second per request, against 189. When generating, the engine scores its whole configured context window to choose which earlier tokens to attend to, however short the conversation is. That work grows with both the window and the number of requests.
Keeping short prompts off the long copy did not help
A slow copy that also takes half of the ordinary traffic looked like a problem. We added a Ramjet setting that decides whether short prompts may use the lane (RJ_ROUTE_LONG_PROMPT_SHORT, ramjet#307). We compared three policies on the same engines:
| Short-prompt policy | 16 developers | 32 developers | 64 developers |
|---|---|---|---|
| Shared: short prompts may use the long copy | 123.6 | 181.8 | 230.7 |
| Exclusive: short prompts never use it while the other copy is up | 109.1 | 143.3 | 147.6 |
| Avoid: short prompts use it only once the other copy is busier by a margin | 117.1 | 170.8 | 211.6 |
| Reference: two 262k copies, no lane | 131.4 | n/a | 276.2 |
Shared won at every load, including on latency. Exclusive put every short session on one copy, which halved the server's capacity and stretched the median turn at 16 developers from 2.2 to 2.9 seconds. Avoid failed for a different reason: it moved short prompts by load alone and ignored where each session's prompt was cached, so only 53-60% of turns stayed on their session's copy.
On a two-copy server, the long copy's spare capacity is worth more than isolating short traffic from it. Offering the full one-million-token context cost 6% of throughput at 16 developers and about 16% at 64, compared with two 262k copies. Exclusive lanes make more sense in a larger fleet, where one long copy sits beside several fast ones. A version of avoid that respects where a session is already cached is the obvious next step.
A replica that answered health checks but no requests
During this work, one copy came up, reported itself ready, and then never ran a single request. SGLang's own start-up check request timed out every ten minutes for an hour. Its /health and /v1/models endpoints answered normally the whole time, so a load balancer checking those endpoints would have kept sending it half the traffic.
The cause was outside the GPU. SGLang's web server passes its internal request channels to its tokenizer worker processes through a shared-memory file named after its own process ID (/dev/shm/multi_tokenizer_args_<pid>). Both copies ran in containers that shared the host's /dev/shm. Each container numbers its processes from scratch, so the two servers often got the same process ID. When that happened, the second copy overwrote the first copy's file, and the first copy's workers sent requests into channels that did not exist in their container. Giving each container its own /dev/shm fixed it.
Ramjet already had a check that would catch this: a one-token generation sent to each engine alongside the readiness probe. It ran only for data-parallel ranks of a single engine, the case it was built for on GLM-5.3. We extended it to every upstream (RJ_UPSTREAM_RANK_PROBE=all, ramjet#308) and tested it by freezing one copy's scheduler processes, which reproduces the same state: the web server answers, and nothing generates.
| Generation probe | Ramjet's view of the frozen copy | 12 requests through Ramjet |
|---|---|---|
Data-parallel ranks only (on) | healthy | 6 succeeded, 6 timed out after 30 seconds |
Every upstream (all) | fenced | 12 succeeded in 2 seconds in total |
A recent real completion still overrides a probe timeout, so a copy that is busy serving traffic is not fenced by mistake.
How the engine side was tuned
Routing works on top of engines that were tuned separately. In the order they mattered:
- Dense layers in BF16. The checkpoint stores its non-expert weights in FP8 with 32×32 scaling blocks. On H200, SGLang runs those through a slow general-purpose kernel, which took half of all GPU time while generating. Converting them to BF16 once at load is exact, and took single-request generation from 234 to 392 tokens a second.
- DSpark draft decoding, at a 262k context so its buffers fit: about three times faster for a single request.
- Two copies instead of one engine across eight GPUs: 23% more agent turns at 16 developers and 59% more at 64.
- CPU pinned, memory not. Pinning each GPU's process to its nearest CPUs mattered (5% at 64 developers). But SGLang also confines that process's memory to the same NUMA node, the 110 GB slice of host memory attached to those CPUs (NUMA stands for non-uniform memory access). The 189 GB Engram table then did not fit, and the kernel killed the process. A small change made the memory binding a preference instead.
None of this changed answer quality that we could measure. GSM8K accuracy was 94.8% against 95.2% on the stock engine, inside the run-to-run noise, and Ramjet's agent protocol checks passed in both.
Ramjet is open source, and its experiment journal holds our raw benchmark notes. For the model itself, see DeepSeek V4.1 Flash: Why Its Encoder, Engram and KV Cache Matter. For the same routing on a model whose cache does run out, see Serving Full GLM-5.3 to 48 Coding Agents From One 8×H200 Server.
Measured on one server with 8× H200 141 GB GPUs, running deepseek-ai/DeepSeek-V4.1-Flash at revision 2cba9e42 as two four-GPU copies. The engine was SGLang 0.5.21 with the BF16 dense-layer and NUMA changes described above and DSpark with a 5-token draft. The router was Ramjet 0.7.0, plus the long-prompt and probe changes described here. Each benchmark cell ran for five minutes after a one-minute warm-up with fresh prompt text. The routing comparison ran with NUMA pinning off, so its absolute numbers sit about 5% below the headline figures. These are serving measurements for this machine and workload, not model-quality claims.