If vLLM already solved LLM serving, why did SGLang appear?
At its heart, high-performance computing (HPC) architecture is about teamwork. It’s the art of designing many separate computers to work together seamlessly as one powerful machine. While a single server eventually hits a limit, HPC architecture helps us break through that ceiling. It combines many nodes with a lightning-fast network, high-speed storage, and smart software to keep everything running smoothly until the job is done.
The most important thing to remember is that an HPC cluster has six layers, and your overall performance is only as good as your weakest link. Even the fastest GPUs in the world will be held back by a slow network or a sluggish file system. Most clusters that don’t live up to expectations usually have bottlenecks in the network, storage, or scheduling layers.

The short answer
| Layer | What it is | What it decides | What to specify |
| 1. Compute | GPUs and CPUs, with NVLink and NVSwitch inside the node | The raw ceiling of one node | GPU model, memory per GPU, GPUs per node |
| 2. Fabric | The network between nodes: InfiniBand or RoCE v2 Ethernet | Multi-node throughput, and therefore almost everything | Speed per port, and the oversubscription ratio |
| 3. Storage | A parallel file system: Lustre, IBM Storage Scale, WEKA, VAST, DDN | Whether the GPUs stay fed through reads and checkpoints | Aggregate throughput in GB/s, not just capacity |
| 4. Software | MPI, NCCL or RCCL, CUDA or ROCm, containers | Whether your code can reach the hardware | Driver and toolkit versions, container runtime |
| 5. Scheduler | Slurm, or Kubernetes | Where your ranks physically land | Topology awareness, queue policy, preemption |
| 6. Operations | Monitoring, node replacement, escalation | Whether a multi-week run finishes | Replacement SLA in hours, proactive detection rate |
Let’s walk through each layer in order to see how they build upon one another.
Layer 1. Compute: Where the Power Begins
Think of an HPC node as a high-powered server filled with “accelerators.” For modern AI and scientific research, these accelerators are almost always GPUs (Graphics Processing Units).
The reason is arithmetic throughput. A CPU has tens of cores optimised for branching, latency-sensitive work. A GPU has tens of thousands of simpler cores optimised for doing the same operation across a very large array at once. Training a neural network, simulating a fluid, or pricing a derivatives book are all that second shape of problem.
When looking at GPUs, two specific features really stand out.

Memory bandwidth determines your speed. For tasks like generating text with AI, the bottleneck is often how quickly the chip can read data from memory. For example, while an H200 and H100 have similar raw power, the H200’s faster memory allows it to generate tokens much more quickly.
Memory capacity determines how much hardware you need. This is a huge cost factor that is easy to overlook. If a model is too big for one chip, you have to buy more. Using chips with larger memory can significantly reduce the total number of GPUs you need to manage the same workload.

A 405-billion-parameter model at FP8 needs seven 80GB H100s to hold. On 288GB parts it needs two. Same model, a third of the hardware.
CPUs still matter, as the host. They run the operating system, feed data to the GPUs, and handle everything that does not vectorise. Underspecify them and you build a cluster where expensive accelerators wait on data loaders. ARM-based host CPUs, including NVIDIA’s Grace line, offer far higher CPU-to-GPU bandwidth than PCIe, which matters when you are offloading caches or LoRA weights to host memory.
Layer 2. Inside the Node: Getting GPUs to Talk
Inside a single server, GPUs are linked together so they can share information almost instantly. NVIDIA uses NVLink and NVSwitch to create this “all-to-all” connection, while AMD uses Infinity Fabric. This internal speed is why we treat the node as the basic building block of any cluster.
- NVLink is the point-to-point link between GPUs, running at 900 GB/s on Hopper and 1,800 GB/s on Blackwell.
- NVSwitch sits on top of NVLink and provides all-to-all communication, so every GPU in the node can talk to every other at full rate.
- AMD Instinct uses Infinity Fabric for the equivalent job.
This is why the node is the natural unit of an HPC architecture, and it explains the oldest distinction in the field.
| Model | What it means | Where you see it |
| Shared memory | Every processor addresses one common pool of memory. Simple to program, limited by how much you can physically attach | Inside a node. Programmed with OpenMP |
| Distributed memory | Each node has its own memory. Nodes exchange data by sending messages explicitly | Across nodes. Programmed with MPI |
| Hybrid | Shared memory within the node, distributed memory across nodes | Every real cluster today. OpenMP or CUDA inside, MPI or NCCL across |
Almost every modern system uses a “hybrid” model. This means you write code that treats the GPUs inside a node as a single, tightly-knit unit, while being very careful about how you send data over the network between different nodes.
The word “expensive” is doing real work in that sentence.
Layer 3. The Fabric: Connecting the Dots
When you move from communicating inside a node to communicating between nodes, speed takes a significant hit. Bridging this gap is the real challenge of HPC architecture.

Inside a Blackwell node, GPUs talk at 1,800 GB/s. The best node-to-node link available drops that to roughly 100 GB/s. That is an order of magnitude, and designing around it is most of what makes HPC architecture a discipline rather than a shopping list.
Why speed matters: In AI training, nodes constantly need to pause and “agree” on the data they’ve processed (a process called an all-reduce). If the network is slow, your expensive GPUs will sit idle waiting for updates. On a high-speed InfiniBand network, this might take 3 seconds; on standard Ethernet, it could take 20 seconds or more. This idle time is a common reason for poor cluster performance.
A collective runs at the speed of its slowest path.
Work the arithmetic. A 70-billion-parameter model in bf16 holds 140 GB of gradients. A ring all-reduce moves roughly twice that across each node’s uplink, so about 280 GB per node per step.
- On an 800G InfiniBand fabric, roughly 100 GB/s, that exchange takes under 3 seconds.
- On standard 100 gigabit Ethernet, roughly 12.5 GB/s, it takes over 20 seconds.
Frameworks overlap communication with computation, so you never pay the full number. But you cannot overlap bandwidth you do not have. On the wrong fabric the GPUs sit idle, and your dashboard reports it as low GPU utilisation rather than as a network problem. That is a common misdiagnosis in multi-node training.
The interconnect options
| Interconnect | What it is | Strengths | Watch for |
| NVIDIA Quantum InfiniBand | Purpose-built HPC fabric, now at 800G | Congestion control managed in the fabric, predictable at scale, assumed by NVIDIA reference architectures | Single supplier, higher cost per port |
| RoCE v2 over Ethernet | RDMA over Converged Ethernet, at 400G and 800G | Standard switches, multi-vendor, no single supplier, works well with AMD Instinct | Needs careful DCQCN and PFC tuning to match InfiniBand behaviour |
| Omni-Path | Originally Intel, now Cornelis Networks | An alternative HPC fabric with an installed base in traditional HPC | Smaller ecosystem in AI clusters |
| Standard Ethernet, no RDMA | Ordinary TCP/IP networking | Cheap, universal | Not viable for tightly coupled multi-node training |
Whether you choose InfiniBand or RoCE v2 Ethernet depends on your needs. InfiniBand is the classic, high-performance choice often validated by NVIDIA, while open Ethernet offers more flexibility and vendor independence. The most important thing is that your provider gives you a choice that fits your workload.
Topology, and the ratio nobody volunteers
Switches are arranged in a topology. Two names dominate.
- Fat-tree, also called leaf-spine. Nodes connect to leaf switches, leaf switches connect upward to spine switches. Bandwidth increases as you move up the tree, so the upper levels can carry the traffic of everything beneath them. This is what almost every AI cluster uses.
- Dragonfly. Switches are grouped, groups are richly interconnected, and most traffic stays local. It reduces cable count and cost at very large scale, and appears more in national laboratory systems than in commercial AI clusters.
Whichever topology you get, one number decides what you actually receive.

- Non-blocking, or 1:1. The uplink capacity leaving a leaf switch matches the capacity of the ports beneath it. Every node can transmit at full rate simultaneously.
- Oversubscribed, often 3:1. The uplink carries a third of the capacity beneath it. This is a sensible, cost-saving design for web servers, which rarely all transmit at once. In an all-reduce every node transmits at once, so you receive a third of the fabric you thought you bought.
A “non-blocking” (1:1) network ensures every node can talk at full speed at the same time, which is exactly what you want for HPC.
Layer 4. Storage: Fuel for Your GPUs
To keep thousands of processes running at once, you need a parallel file system. This allows many servers to read and write data simultaneously without getting in each other’s way.
| System | Type | Typically used for |
| Lustre | Open source, the traditional HPC standard | National labs, academic clusters, and managed as a service by most large clouds |
| IBM Storage Scale (GPFS) | Commercial, POSIX parallel file system | Enterprise HPC, mixed workloads |
| WEKA | Commercial, NVMe-first | AI clusters where checkpoint and small-file performance matter |
| VAST Data | Commercial, disaggregated shared-everything | Large AI estates with mixed hot and warm data |
| DDN | Commercial, including EXAScaler | Large training clusters, long-established in HPC |
| Object storage | S3-compatible | Datasets and archives. Not a substitute for a parallel file system at cluster scale |
Fast storage is vital because of “checkpointing.” Every hour or so, a cluster saves its entire state so it doesn’t lose work if something fails. A single checkpoint can be 1 TB or more. If your storage is slow, your entire cluster stays idle while it saves. Over a week-long run, the difference between slow and fast storage can save you hours of expensive compute time.
A 70-billion-parameter model does not checkpoint 140 GB. It checkpoints the full training state: bf16 weights, an fp32 master copy, and two Adam optimizer moments. That comes to roughly 12 to 14 bytes per parameter, so about 1 TB per checkpoint. Checkpoint hourly, which is normal at that scale, and a seven-day run writes it 168 times, with the entire cluster idle for each write.

That is the same GPUs, the same job, and a difference of four hours of paid cluster time. Storage throughput is a compute cost line, not an IT cost line.
If you want to dive deeper into how different types of storage compare, our guide on object vs block vs file storage is a great place to start.
Layer 5. The Software Stack: Making It All Work
Hardware is just expensive decoration without the right software. The “stack” includes the languages, libraries, and frameworks that allow your code to actually talk to the chips and the network.
| Layer | Options | What it does |
| Parallel programming model | MPI, OpenMP, CUDA, ROCm, SYCL | MPI passes messages across nodes. OpenMP parallelises within a node. CUDA and ROCm target NVIDIA and AMD GPUs |
| Collective libraries | NCCL, RCCL, MPI collectives | Implement all-reduce, all-gather and broadcast efficiently over the fabric. This is where topology awareness pays off |
| Frameworks | PyTorch, JAX, DeepSpeed, Megatron-LM, vLLM, SGLang | What you actually write against. Most call NCCL underneath without you seeing it |
| Containers | Docker, Apptainer (formerly Singularity), Enroot with Pyxis | Package the environment so a job runs the same on every node. Apptainer and Enroot are preferred on HPC because they run unprivileged under a batch scheduler |
| Languages | C, C++, Fortran, Python | Fortran and C still dominate legacy scientific codes. Python dominates AI, with the heavy lifting in compiled kernels underneath |
Many clusters that suddenly stop working have run into a version mismatch.
Layer 6. The Scheduler: Directing Traffic
Slurm is the standard tool used to manage HPC jobs. You tell it what you need, and it finds the right physical nodes for your job as they become available. For more dynamic workloads, many teams are also using Kubernetes.
By default, Slurm does not know how those nodes are wired. You tell it, through a file called topology.conf, which describes which nodes sit beneath which switch.
It is vital that your scheduler “understands” your network layout. If it places parts of a job on nodes that are physically far apart in the network, your communication will slow down significantly without any obvious error message in the logs.
Three ways a cluster is handed to you:
- Pure bare metal. Root access and your own kernel. Take it if you have a platform team and strong opinions.
- Bare metal with Slurm. Batch queueing, fair-share policy and checkpoint recovery, configured before handover. The default for classical training runs.
- Bare metal with Kubernetes. Container orchestration, better suited to workloads that mix training with serving or that scale dynamically. See Kubernetes worker nodes explained.
Slurm and Kubernetes are not rivals so much as different assumptions. Slurm assumes a queue of finite jobs competing for a fixed pool. Kubernetes assumes long-running services that scale with demand. Large sites increasingly run both.
Layer 7. Operations: The Secret to Finishing
This layer doesn’t usually show up in technical diagrams, but it’s often the most important. Computers at this scale don’t always fail cleanly—sometimes a chip just slows down slightly, which can drag down the performance of your entire project.
GPUs rarely fail cleanly. They throttle under thermal pressure, log correctable ECC errors, or slow just enough to lag their rank. In a synchronous training job, one slow rank sets the pace for the entire cluster.
Then there is outright node loss, and the arithmetic is unforgiving.

A single node might be reliable. Five hundred of them running for a week are not. Two numbers cap your exposure:
- How long a failed node stays in your allocation. A four-hour replacement instead of thirty minutes loses three and a half hours of the whole cluster, not of one node.
- What share of incidents get caught before you report them. The difference between a swap that happens while you sleep and one that starts when you file a ticket.
The goal of good operations is goodput: the actual amount of useful work you keep. This involves proactive monitoring and quick hardware replacement. You can learn more about this in our article on GPU monitoring.
Putting It All Together
These layers are all connected. For example, if your storage is slow, you might decide to save checkpoints less often. But if a node fails, you lose more work between those saves. You end up spending more time redoing calculations, even if your GPU graphs look “busy.”
- Your file system delivers 10 GB/s instead of 100, so each checkpoint takes 100 seconds instead of 10.
- To reduce the number of checkpoints, you lengthen the interval from one hour to four.
- A node fails on hour three of an interval, as it eventually will on a large cluster.
- You restart from a checkpoint that is three hours old, on top of the node replacement time.
- Your GPU utilisation graph looks fine throughout, because the GPUs really were busy. They were just busy redoing work.
Nothing in that sequence is a GPU problem, and no GPU benchmark would have found it.
Benchmarking: how you know the architecture works
| Benchmark | What it measures | When to use it |
| HPL / LINPACK | Dense linear algebra across the whole machine. The basis of the TOP500 list | Proving a cluster works end to end. A TOP500 entry means the fabric has been exercised, not just described |
| HPCG | Sparse workloads, a harder and more realistic memory-bound test | Where LINPACK flatters the system |
| NCCL tests | All-reduce and all-gather bandwidth at your actual node count | The single most useful acceptance test for an AI cluster |
| MLPerf Training and Inference | Standardised model workloads, audited | Comparing platforms on work that resembles yours |
| IO500 | Storage performance under HPC access patterns | Validating layer 3 before you sign |
Before you commit to a cluster, always run real-world tests at full scale. A network that works for 8 nodes might struggle at 64. Testing early ensures you get the performance you’re paying for.
For monitoring in production, Prometheus with Grafana and NVIDIA DCGM Exporter has largely replaced the older Ganglia and Nagios pairing, though both still run at established sites.
Build or Rent?
| Build your own | HPC as a service | |
| Capital | Large upfront, depreciating over 3 to 5 years | Operating expense, matched to the project |
| Time to first job | 9 to 18 months including power, space and commissioning | Weeks |
| Silicon refresh | You own the obsolescence | The provider does |
| Utilisation risk | You pay for the trough | Committed base plus on-demand peaks |
| Expertise needed | Facilities, networking, storage, Slurm, 24×7 operations | Your workload, and a good contract |
| Best when | Continuous high utilisation, extreme data-gravity or security constraints | Almost everything else |
Should you build your own or use a service? If you can keep a cluster busy more than 70% of the time for three years, building might be for you. Otherwise, HPC as a Service or using an AI neocloud is often much more cost-effective.
How Neysa Can Help
Neysa provides dedicated HPC clusters based in India. We are proud to be the only India-headquartered provider highly rated in the SemiAnalysis ClusterMAX evaluation, which tests real-world performance on rented clusters.
Mapped to the six layers above, every one is a choice rather than a default:
- Compute. NVIDIA and AMD, current generation. Around 2,000 GPUs live today, with 8,192 air-cooled AMD Instinct MI350X landing as a single-site cluster and 9,216 air-cooled NVIDIA B300 staged between December 2026 and March 2027, on a path past 25,000 GPUs by 2027.
- Fabric. NVIDIA Quantum-X800 InfiniBand at 800G where you want the validated reference architecture, or non-blocking open-standards RoCE v2 Ethernet at 400G and 800G where you want vendor independence. Both offered non-blocking.
- Storage. WEKA, VAST or DDN, open-source software-defined storage, or your own licences running on Neysa hardware.
- Software and scheduler. Pure bare metal, bare metal with Slurm pre-configured for batch and checkpoint recovery, or bare metal with Kubernetes.
- Operations. A 99.9% uptime SLA held 100% for six months running, 30-minute first response on Severity-1 against a four-hour resolution target, a GPU-replacement SLA written into the contract with instances restored in 30 to 60 minutes, and roughly one incident in three caught by monitoring before the customer reports it.
- Security and Compliance. We maintain the highest standards, including ISO and SOC 2 certifications. You can find more details on our infrastructure security page.
Our clients, like ITQ Technologies, TIFIN, and Innoviti, have seen significant cost savings and performance improvements by using our tailored HPC solutions.
Interested in seeing how we compare to others? Take a look at our list of the Top 10 HPC Cloud Providers in India.
See predictable GPU cloud pricing for your workload. Talk to an AI Expert about deploying on Neysa Velocis.










