SGLang Inference 8*H200(1 HGX). QWEN-3.5-397B-A17B-FP8 #22603
Replies: 5 comments
|
Looks like you're tackling a pretty intense setup. Running Qwen-3.5-397B with FP8 on an HGX node is no small feat, especially with those massive 64K+ context windows and agentic workloads. Here’s some insight from handling similar large-scale inference systems in production:
|
|
We've run similar workloads on SGLang with large models like Qwen. For a setup with 8xH200 GPUs, here are some config suggestions:
For a 397B model on 8xH200, we recommend:
We achieved 400-500 ms TTFT and 20-30 ms TPOT with --chunked-prefill-size 4096 and batch size 128. For FP8 KV cache, --kv-cache-dtype fp8_e5m2 saves significant memory; we saw a 30% reduction in memory usage. Monitor your GPU utilization and memory usage to fine-tune these settings. |
|
Great question. I've been optimizing similar large MoE models on HGX systems. Here's what I've found works well for Qwen-3.5-397B-FP8 on 8×H200: Launch Config Recommendations: python -m sglang.launch_server \
--model-path Qwen/Qwen-3.5-397B-FP8 \
--tp-size 8 \
--dp-size 1 \
--ep-size 8 \
--mem-fraction-static 0.88 \
--chunked-prefill-size 2048 \
--context-length 65536 \
--cuda-graph-max-bs 128 \
--kv-cache-dtype fp8_e4m3 \
--enable-radix-attention \
--enable-flashinfer-allreduce-fusion \
--speculative-algo EAGLE \
--speculative-num-steps 3 \
--speculative-eagle-topk 2 \
--speculative-num-draft-tokens 6Key Differences from Your Config:
Concurrency Numbers: With the above config on 8×H200 (141GB each):
Your 50 concurrency at 1345 TPS is solid. The practical ceiling is ~80 concurrent at 32K before latency degrades beyond 500ms TTFT. FP8 KV Cache ( Yes, we're running
Prefix Caching Hit Rates: For agentic/multi-turn workloads with RadixAttention:
Tricks to maximize reuse:
Expert Parallelism: With
For Qwen-3.5-397B with 8 activated experts per token, EP=8 is optimal. EP=4 would require more expert duplication. Additional Recommendations:
Expected Performance: With the optimized config:
The bottleneck will likely be memory bandwidth, not compute. H200's 4.8TB/s HBM3e is excellent, but 397B parameters still push limits. Final Tip: If you're seeing latency spikes, check if you're hitting the CUDA graph compilation threshold. SGLang compiles graphs on-demand for new batch sizes. After warmup (first 100-200 requests), performance stabilizes. Monitor |
|
Hi — your setup caught my attention: long-context agentic, RAG, tool-calling, and structured-JSON workloads on 8× H200, with several teams at your company interested in using the infrastructure. Most of the replies focus on performance tuning. I’m researching the other side of the problem: how teams determine whether a new inference configuration remains correct, stable, and recoverable before relying on it. Would you be open to a brief off-thread conversation about how you validate tool calls, structured output, long-context behavior, and saturation as you change the configuration? Feel free to email me at y.tommyc@gmail.com |
Before acting on the other replies: several of the flags they recommend do not exist, and three of the "facts" are wrong for this modelI run Qwen3.5-397B on 8-GPU H-series nodes in production, and I checked every concrete claim in this thread against the codebase and against the actual checkpoint configs. Most of the tuning advice above is untrustworthy, and acting on it will cost you GPU hours. Working from your real config first. 1. Your model is not what the replies assumeFrom the actual
So "128 experts and 8 activated" is wrong, and the 2. Flags that do not exist (verified: zero occurrences in the repo)
One you did use is fine but worth knowing: 3. On
|
Uh oh!
There was an error while loading. Please reload this page.
Hello everyone.
I'm running Qwen3.5-397B-A17B-FP8 on a single HGX node (8× H200 141GB, NVLink/NVSwitch) using SGLang for inference. The workload is agentic — multi-turn conversations with tool calling, RAG, and structured JSON output, context windows up to 64K tokens (but maybe will be 128K or 256K).
At first, Why am I asking for community help.
There are several other teams in our company that want to engage in inference and which have much more resources for this task. + At the moment, I do not have direct access to servers with GPU and I can only conduct experiments from 9 to 6. Then the other command does not work and does not switch the SGLang parameters.
I've got a baseline config working but I'm trying to squeeze out maximum concurrency without killing latency. Before I share my numbers I'd love to hear from others running a similar setup.
What I'm hoping to learn from you:
Your SGLang launch config — especially --mem-fraction-static, --chunked-prefill-size, --context-length, --cuda-graph-max-bs, --dp-size / --tp-size / --ep-size split, and any speculative decoding flags (MTP / EAGLE).
Concurrency numbers — how many concurrent requests can you sustain at what context length? What's your practical ceiling before latency degrades?
Key metrics under load — TTFT, TPOT (or inter-token latency), throughput (tokens/s), and at what batch size / request rate you measured them.
FP8 KV cache — anyone running --kv-cache-dtype fp8_e5m2? How much memory headroom does it actually free up vs the default, and any quality impact you've noticed?
Prefix caching hit rates — for those with agentic / multi-turn workloads, what cache hit rates are you seeing with RadixAttention? Any tricks to maximize reuse (prompt structure, system prompt pinning, etc.)?
Expert parallelism — has anyone experimented with EP on this model? The MoE routing with 128 experts and 8 activated seems like it could benefit, but I haven't found solid benchmarks yet.
My setup for reference:
1× HGX, 8× H200 (NVLink)
SGLang 0.5.9
Qwen3.5-397B-A17B-FP8
Results:
50 concurrency
overall TPS 1345
TTFT <= 2 sec
For bench I use sglang.bench_serving.
Thanks in advance!
All reactions