GLM-5.3-Flash on 8x B200: dual TP4/EP4 serving profile with long-context measurements #37153
Replies: 2 comments
|
Update, September 27: a new same-day run on the same 8x B200 setup found four things.
The run also covered 16K to 1M input at 1 to 16 concurrent requests: 744 of 744 requests succeeded, three repetitions per cell. Every result ran with the local patch and a tool-call parser overlay. The TP8 run came after the main grid and used different prompts. The output rate is for one burst of requests, not sustained load. The original post said the raw traces were unpublished; the request-level data is now in the repository, and |
|
Revision note (September 27): the post now opens with the 1M out-of-memory finding from the comment above, and its last paragraph notes that the benchmark harness and request-level rows have been published since September 3. No measured values changed. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
I measured an SGLang serving profile for
zai-org/GLM-5.3-Flashon one 8x NVIDIA B200 node. It uses two independent TP4/EP4 replicas, FP8 KV cache, TRT-LLM DSA backends, the FlashInfer TRT-LLM MoE runner, static NEXTN 5-1-6, and session-affine routing.Update, September 27: on the pinned image, one server prefilling two 1M-token prompts together ran out of GPU memory in the sparse-attention indexer. #40854 fixes this upstream, and the repository now publishes an equivalent local patch. The results below predate the patch. The September run's findings are in this comment.
The pinned software stack, engine commands, proxy configuration, client setup, methods, and results are documented here: https://github.com/stewtong/b200-glm53
Validation observations:
reasoning_effort: low. The conditions used different output limits and are reported separately in the repository.The serving table and cache study used three repetitions per condition. Since September 3, the repository also publishes the benchmark harness and the request-level rows for both. The repository documents the tested boundaries and excludes unsupported aggregate throughput, decode TPOT, route-latency, and causal claims.
All reactions