HelixML

Serving Full GLM-5.3 to 48 Coding Agents From One 8×H200 Server

Sep 29, 2026

The full 753-billion-parameter GLM-5.3 fits on one 8×H200 server. Splitting its attention eight ways, routing each agent session back to its own GPU with Ramjet and adding a host-memory cache took the server from 30 to about 100 agent turns a minute, enough for 32 to 48 continuously working coding agents.

The full GLM-5.3 checkpoint is 755.7 GB of FP8 weights. An 8×H200 server has 1,128 GB of GPU memory, so the model fits on one machine with room left over for the cache that makes agent traffic cheap. How much room, and how well it is used, decides how many developers that server can carry.

We served GLM-5.3 on one 8×H200 server to a simulated team of developers, each driving a coding agent that never stops working. The straightforward layout, with every layer split across all eight GPUs, fell over at 64 developers. It managed 29.6 agent turns a minute, the median turn waited 94 seconds for its first token, and only 59% of prompt tokens came from cache. The final setup handles 32 developers at a 1.3-second median wait and 48 at 108.5 turns a minute, with 91-93% of every prompt served from cache and no failed requests.

Three changes got it there. Two of them are engine settings. The third is Ramjet, the open-source load balancer we run in front of our inference engines.

The workload: agents that re-send their whole conversation

A coding agent sends its entire conversation on every turn: the harness's system prompt and tool list, the repository's instructions, then every message, tool call and tool result so far. The prompt only grows until the agent compacts it. Most turns add a tool result of a few thousand tokens to a prompt of tens of thousands.

So the engine's work is dominated by reading prompts, and nearly all of each prompt was already read on the previous turn. If the engine still holds that earlier work in GPU memory (the key-value cache, or KV cache), a turn costs only its new tokens. If it does not, the engine reads the whole conversation again. For GLM-5.3 on this server that took 11 to 16 seconds for an 86,000-token conversation.

Our benchmark (agent_swarm_bench.py) simulates that team. Each developer uses one of three agent harnesses with an 18,000-token system prompt and works in one of six repositories. Tool calls take about two seconds, humans rarely pause, and when a conversation passes 110,000 tokens the agent compacts it. Prompts averaged 34,000-43,000 tokens and reached 115,000. Each cell ran for 5 to 15 minutes after a warm-up. The prompts are synthetic filler, fresh for every run.

A "developer" here is an agent that is busy almost all the time. That is a much heavier load than a person, whose agent sits idle while they read, think, review and go to meetings.

Splitting the model eight ways left too little cache

The usual way to run a model this size on eight GPUs is to split every layer across all of them (tensor parallelism, TP8). Each GPU then holds about 94 GB of weights. The catch is in GLM-5.3's attention. Its compressed attention cache cannot be divided between GPUs the way ordinary attention can, so every GPU keeps a full copy of it. With eight identical copies, the whole server held 353,000 tokens of conversation, about ten agents' worth.

At 16 developers that was enough: 60.7 turns a minute and 89.5% of prompt tokens from cache. At 64 it was not. Conversations evicted each other, the engine re-read them from scratch, and throughput halved to 29.6 turns a minute.

Eight data-parallel ranks, eight private caches

SGLang can run the attention layers data-parallel instead (DP attention). Each GPU becomes a rank that handles its own requests with its own cache, while the model's experts stay spread across all eight GPUs. Combined with an 8-bit cache format (FP8 KV), that took the server from 353,000 cached tokens to 1.55 million, 193,000 per rank.

Coding agents send requests to ramjet, which lists each of the engine's eight data-parallel ranks as its own upstream and sends every turn of a session back to the rank holding its prompt. Each rank has its own GPU KV cache and a 32 GB host memory tier.

The new problem is that a rank can only reuse a prompt it has read itself. SGLang assigns incoming requests to ranks in turn, so an agent's consecutive turns land on different GPUs, and each one finds a cold cache. With 16 developers the cache hit rate fell from 89.5% to 65.6%, and throughput fell to 43.0 turns a minute, below the TP8 layout we were trying to beat.

Ramjet sends each session back to its rank

Ramjet already sends each conversation back to the replica that holds its prompt. It does this by fingerprinting the start of each request and remembering where it went, then weighing that against how busy each replica is. For this server we taught it to treat each DP rank as a separate replica. It lists the same engine URL eight times and pins each request to a rank with SGLang's routed_dp_rank field (RJ_UPSTREAM_DP_RANKS).

With 16 developers, 94.5% of turns landed on the rank that had served the session before.

At 16 developers: TP8 gave 60.7 turns a minute with 89.5% cached and a 4.4-second p90 wait; eight ranks placed by SGLang gave 43.0, 65.6% and 5.0 seconds; eight ranks placed by ramjet gave 72.7, 92.8% and 1.7 seconds.

Same server, same model and the same simulated team in each row.

That is 69% more turns than the engine's own placement on identical hardware, and 20% more than TP8. The slowest one in ten turns started within 1.7 seconds instead of 5.

Rank-level routing needs one more piece. SGLang keeps answering its health check when a single rank's scheduler has hung, so a load balancer would keep sending that rank's sessions into it. Ramjet v0.7.0 can probe each rank with a one-token generation pinned to it (RJ_UPSTREAM_RANK_PROBE=on), and a recent real completion on the rank counts as proof it is alive, so busy ranks are not fenced. The probe is on in our deployment file. It was not active during the measurements in this post.

Eight ranks also change how Ramjet scores affinity. Its default compares each replica's cached prefix against a fixed cap. Agent harness prompts are longer than that cap, so every rank that had seen the harness scored full marks and session history became a tie-break. The relative setting scores each rank against the warmest rank instead, which keeps sessions sticky across eight or more replicas.

Host memory as a second cache tier

Pinning fixed placement, but at 64 developers the working set outgrew even 1.55 million tokens. Ranks still evicted conversations that came back a minute later: 45.9 turns a minute and 71.8% cached.

SGLang's hierarchical cache (HiCache) copies evicted cache pages to the server's ordinary memory and copies them back when the conversation returns. We gave each rank 32 GB of host memory for it. On an idle rank, an 86,000-token conversation came back from host memory in 0.8 to 1.0 seconds, against 11 to 16 seconds to read it again from scratch.

At 64 developers, adding 32 GB of host memory per rank took the server from 45.9 to 95.1 turns a minute, and from 71.8% to 91.8% of prompt tokens cached.

Both rows use eight ranks and Ramjet routing.

It doubled the overloaded server, from 45.9 to 95.1 turns a minute. Prompt tokens processed rose from 27,700 to 66,700 a second. With 16 developers, when everything already fitted on the GPUs, the tier made no difference.

Before trusting it, we flooded a rank until an 86,000-token prompt was pushed out to host memory. We then sent the prompt back and checked that the model still recalled a fact from it and produced a correctly typed tool call. Every engine configuration with the host tier passed that check before we measured it.

How many developers one server carries

With all three changes in place, we raised the team size until throughput stopped growing.

Turns per minute rise from 73.2 at 16 developers to 98.2 at 32 and 108.5 at 48, then stay flat at 104.6 and 101.7 for 64 and 96. The median and p90 wait for the first token grow from 0.84 and 1.7 seconds at 16 developers to 18.6 and 80 seconds at 96.

The 32- and 48-developer cells and the 104.6 at 64 ran on a second server of the same type. At 16 developers the two servers agreed within 2.5%, but at 64 the second ran 10% faster, so treat differences under 10% at high load as unresolved.

DevelopersTurns / minPrompt tokens / sWait for first token, median / p90Turns starting within 5 s
1673.244,5000.84 s / 1.7 s99.6%
3298.270,9001.27 s / 7.0 s85%
48108.577,0002.77 s / 17.5 s61%
6495.1-104.666,700-73,90011.4 s / 46.1 s26%
96101.765,20018.6 s / 80 s13%

The server tops out at around 100-110 agent turns a minute. Past 48 developers, adding more only makes everyone wait longer.

So one 8×H200 server carries 32 to 48 continuously working coding agents, depending on how long you will let a turn wait:

  • 32 agents if most turns should start within a few seconds. The median wait is 1.3 seconds and 85% of turns start within 5.
  • 48 agents for the most work per server. Throughput is at its peak, but the slowest one in ten turns waits 17.5 seconds.

Those are agents that are almost never idle. Each one at 48 still completes a turn about every 27 seconds. How many people that covers depends on how much of the day their agents spend working. If a typical engineer's agent is active for a third of the working day, 32 to 48 busy agents is roughly 100 to 150 engineers. We have not measured that ratio, and it will vary by team.

Renting the server against paying per token

At 32 agents the server reads 70,900 prompt tokens and writes 452 output tokens a second, with 92.8% of the prompt coming from cache. Priced at list rates, one hour of that traffic would cost about $172 on Claude Opus 5.5 and $110 on OpenAI's GPT-6 Sol, the model behind Codex. Renting an 8×H200 server costs roughly $24-36 an hour.

The calculator below uses the measured token rates at each load level. Set your rental price and how many hours a day your agents run.

Load on the server (measured)
H200 rental $ / GPU-hour$3.50 · $28/h per server
Hours a working day the agents run10 h
Claude Opus 5.5 · $ / 1M tokens
OpenAI GPT-6 Sol (Codex) · $ / 1M tokens
Rented 8x H200 · GLM-5.3
$20,440/mo
$639 per agent · $28/h
Same tokens on Claude Opus 5.5
$37,815/mo
$1,182 per agent · $172/busy hour
Same tokens on GPT-6 Sol
$24,121/mo
$754 per agent · $110/busy hour
At 10 busy hours a working day, 32 agents send 56.2B prompt tokens and receive 358M output tokens a month. A reserved server costs the same whether it is busy or not. The API bill overtakes it once the agents run 5.4 h a day on Opus 5.5 or 8.5 h a day on GPT-6 Sol.

Server throughput is measured: full GLM-5.3 on one 8x H200 server behind ramjet, with simulated coding agents that are busy almost all the time (the 32- and 48-agent points ran on a second server of the same type). The API columns price the same token counts at list rates: cached prompt tokens at the cache-read rate, uncached prompt tokens at the cache-write rate, output at the output rate. They are not a quality comparison. Opus 5.5 and GPT-6 Sol use different tokenizers and may think for more or fewer tokens to do the same task, and an API's cache hit rate depends on your harness. Rental prices vary by provider and commitment; the default is $3.50 per GPU-hour. 22 working days a month; a reserved server is billed for 730 hours.

These are token counts from GLM-5.3 priced at other models' rates, not a like-for-like quality comparison. Opus 5.5 and GPT-6 Sol count tokens differently and may spend more or fewer of them on the same task. Agent traffic re-sends so many prompt tokens that the cached part alone costs about $47 an hour at 32 agents, even at $0.20 per million. That is more than the whole server.

A lower-latency variant

GLM-5.3's experts are spread across the eight GPUs, and each generation step exchanges tokens between them. Our default uses DeepEP for that exchange, following the published H200 recipe. SGLang's simpler all-gather path (GLM_MOE_A2A_BACKEND=none in our deployment file, plus a few related switches listed in its README) uses less memory, which leaves 22% more cache per rank, but costs compute.

DevelopersDeepEP turns / minp90 waitAll-gather turns / minp90 wait
1673.21.7 s62.72.0 s
3298.27.0 s87.64.9 s
48108.517.5 s101.210.5 s
6495.1-104.646.1 s101.330.8 s
96101.780 s106.360 s

DeepEP is clearly faster at light load, where both ran on the same server. From 48 developers up, the all-gather path has the shorter tail: 80% of its turns start within 5 seconds at 48 developers, against 61%. Its 7-12% throughput deficit at 32 and 48 developers is within the spread between our two servers. If your target is a p90 wait under 5 seconds, it carries 32 agents at 4.9 seconds.

What we tried that did not win

A single shared cache with vLLM. vLLM v0.30.0 with decode context parallelism splits the attention cache across GPUs instead of copying it, giving one 4.8-million-token pool that needs no routing at all. It had the tightest wait at 64 developers (p90 18.3 seconds) and the fastest per-token generation at light load, but it saturated earlier: 65.1, 92.3 and 86.7 turns a minute at 16, 64 and 96 developers. We did not test its host-memory offload.

The same feature in SGLang. The patches that add context parallelism to SGLang need a newer version than the v0.5.20 we run, and we abandoned the port after two attempts.

A GPU fabric channel for TP8. Enabling IMEX made no measurable difference to the TP8 layout.

More running requests. Raising the limit from 16 to 32 running requests per rank changed nothing at 48 or 64 developers. The ranks were short of cache, not batch slots.

A higher memory setting with larger prompt chunks. 88% GPU memory with 64,000-token prefill chunks ran out of memory under load. The deployment uses 85% and 32,000.

Running it

The deployment is a single Compose file, deploy/glm53_h200 in the Ramjet repository, with a validator that checks it before start-up. It runs one SGLang v0.5.20 engine across all eight GPUs with:

  • DP attention across 8 ranks and FP8 KV cache;
  • DeepEP expert parallelism across 8 GPUs;
  • speculative decoding with GLM-5.3's multi-token prediction head (EAGLE 1/1/2);
  • 32 GB of host cache per rank;
  • 4 tokenizer processes and a /health endpoint that does not generate, so a busy engine does not fail Ramjet's health checks;

and Ramjet v0.7.0 in front with one upstream per rank, relative affinity and the per-rank probe. Cold starts took 7 to 17 minutes. Keep the compiled-kernel cache on persistent disk so restarts are faster.

The measurements behind every number here, including the rejected configurations, are in the Ramjet experiment journal under 29 September.


This follows our series on GLM-5.3-Flash, the smaller model we serve in production on eight RTX PRO 6000 GPUs, most recently on why its cache ran out of snapshots before tokens. Ramjet is open source under Apache-2.0. Talk to Helix about running GLM-5.3 on your own hardware.