Does sglang support --load_format sharded_state #3197
Replies: 1 comment
Yes —
|
| source | throughput |
|---|---|
| local NVMe | ~1623 MB/s |
| network/shared filesystem | ~88 MB/s |
An ~18× gap. For a 1450 GiB checkpoint that is the difference between ~15 minutes and ~4.7 hours — which brackets your 1.5 h nicely and suggests you're on the slow path. Check what --model-path actually resolves to: a local disk path, or a network mount. If it's a network mount, copy the checkpoint to local NVMe first; that single change dominated everything else in my case, and no loader flag compensates for it.
A quick way to confirm on your box:
# cold-cache read throughput of one shard
sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'
dd if=<a-shard>.safetensors of=/dev/null bs=16M status=progressIf that reports tens of MB/s rather than GB/s, storage is your problem, not the loader.
Second: if you're on shared/network storage, there's a flag built for exactly your case
--weight-loader-prefetch-checkpoints"Prefetch checkpoint files into OS page cache before loading. Each rank prefetches a fraction of the shards, reducing total network I/O on shared filesystems (NFS/Lustre) from Ncheckpoint to 1checkpoint. Recommended for models on network storage. When enabled, multi-threaded safetensors loading is disabled by default to avoid I/O oversubscription with the prefetch threads; set
enable_multithread_load=truein--model-loader-extra-configto keep multi-threaded loading (e.g. on local NVMe where prefetch is a no-op)."
That "N× → 1×" reduction is the key sentence: with TP=4 and no prefetch, every rank reads every shard, so your 4-node job may be pulling the checkpoint 4× over the network. Enable prefetch and re-measure — this is often a bigger win than changing loader format, and it's one flag.
Third: thread count, and one counter-intuitive result
Multithreaded loading is enable_multithread_load (default true) with num_threads (default 8) in --model-loader-extra-config. My strong recommendation: do not assume more threads is faster on shared storage. In a concurrency sweep on a network filesystem I measured single-shard scan throughput around:
n=1 → 94 MB/s · n=2 → 190 MB/s · n=4 → 191 MB/s · n=8 → 20 MB/s · n=16 → 199–378 MB/s
The n=8 collapse is the striking part, and it lands exactly on sglang's default of 8. If your storage behaves anything like mine, the default thread count may be near a pathological point. Try a couple of values, and remember that with --weight-loader-prefetch-checkpoints on, multithreaded loading is turned off by default precisely to avoid I/O oversubscription — so if you enable prefetch, don't also force 16 threads.
What I'd do, in order
- Measure the storage first (cold
ddabove). If it's a network mount, localize the checkpoint. This alone took my multi-hour load to minutes. - If you must stay on network storage, add
--weight-loader-prefetch-checkpointsand sweep the thread count. - If storage is already fast and you're still slow, then
sharded_stateis the right next step — it turns "every rank reads the whole checkpoint" into "every rank reads only its shard", which is the structural win you were asking about. You need a pre-sharded checkpoint, built viaexamples/runtime/engine/save_sharded_state.py, and then:--load-format sharded_state \ --model-loader-extra-config '{"pattern": "model-rank-{rank}-part-{part}.safetensors"}' - Also consider
--load-format fastsafetensors(uses the fastsafetensors iterator; has anenable_gdsoption in extra config) if you'd rather not convert the checkpoint, andrunai_streamerif you're loading from object storage — both target the I/O path rather than the sharding structure.
One caveat on all the numbers above: they're from my own hardware (H20-class nodes, a specific shared filesystem), measured as cold reads, so treat them as ratios and failure modes rather than absolute expectations for your cluster. The qualitative points — mmap-lazy loading makes progress bars misleading, network storage is the dominant term, per-rank full-checkpoint reads multiply I/O, and the default thread count can be pessimal — should transfer.
If you can share (a) where --model-path physically lives, (b) the cold dd number, and (c) whether you're currently seeing N× network reads, I can tell you which of these four is your actual 1.5 hours.
Uh oh!
There was an error while loading. Please reload this page.
I am currently deploying DeepSeek R1 bf16 with 4 nodes, but the model's loading time is extremely slow, taking approximately 1.5 hours.
Does SGLang support the --load_format sharded_state option, similar to the VLLM framework?
All reactions