logo
Infrastructure

A Guide to HPC Architecture: How Computing Clusters Actually Work


16 mins.

Table of Content

About the author

Divesh Sood Avatar

Head of Product Marketing

At its heart, high-performance computing (HPC) architecture is about teamwork. It’s the art of designing many separate computers to work together seamlessly as one powerful machine. While a single server eventually hits a limit, HPC architecture helps us break through that ceiling. It combines many nodes with a lightning-fast network, high-speed storage, and smart software to keep everything running smoothly until the job is done.

The most important thing to remember is that an HPC cluster has six layers, and your overall performance is only as good as your weakest link. Even the fastest GPUs in the world will be held back by a slow network or a sluggish file system. Most clusters that don’t live up to expectations usually have bottlenecks in the network, storage, or scheduling layers.

The short answer

LayerWhat it isWhat it decidesWhat to specify
1. ComputeGPUs and CPUs, with NVLink and NVSwitch inside the nodeThe raw ceiling of one nodeGPU model, memory per GPU, GPUs per node
2. FabricThe network between nodes: InfiniBand or RoCE v2 EthernetMulti-node throughput, and therefore almost everythingSpeed per port, and the oversubscription ratio
3. StorageA parallel file system: Lustre, IBM Storage Scale, WEKA, VAST, DDNWhether the GPUs stay fed through reads and checkpointsAggregate throughput in GB/s, not just capacity
4. SoftwareMPI, NCCL or RCCL, CUDA or ROCm, containersWhether your code can reach the hardwareDriver and toolkit versions, container runtime
5. SchedulerSlurm, or KubernetesWhere your ranks physically landTopology awareness, queue policy, preemption
6. OperationsMonitoring, node replacement, escalationWhether a multi-week run finishesReplacement SLA in hours, proactive detection rate

Let’s walk through each layer in order to see how they build upon one another.

Layer 1. Compute: Where the Power Begins

Think of an HPC node as a high-powered server filled with “accelerators.” For modern AI and scientific research, these accelerators are almost always GPUs (Graphics Processing Units).

The reason is arithmetic throughput. A CPU has tens of cores optimised for branching, latency-sensitive work. A GPU has tens of thousands of simpler cores optimised for doing the same operation across a very large array at once. Training a neural network, simulating a fluid, or pricing a derivatives book are all that second shape of problem.

When looking at GPUs, two specific features really stand out.

Memory bandwidth determines your speed. For tasks like generating text with AI, the bottleneck is often how quickly the chip can read data from memory. For example, while an H200 and H100 have similar raw power, the H200’s faster memory allows it to generate tokens much more quickly.

Memory capacity determines how much hardware you need. This is a huge cost factor that is easy to overlook. If a model is too big for one chip, you have to buy more. Using chips with larger memory can significantly reduce the total number of GPUs you need to manage the same workload.

A 405-billion-parameter model at FP8 needs seven 80GB H100s to hold. On 288GB parts it needs two. Same model, a third of the hardware.

CPUs still matter, as the host. They run the operating system, feed data to the GPUs, and handle everything that does not vectorise. Underspecify them and you build a cluster where expensive accelerators wait on data loaders. ARM-based host CPUs, including NVIDIA’s Grace line, offer far higher CPU-to-GPU bandwidth than PCIe, which matters when you are offloading caches or LoRA weights to host memory.

Layer 2. Inside the Node: Getting GPUs to Talk

Inside a single server, GPUs are linked together so they can share information almost instantly. NVIDIA uses NVLink and NVSwitch to create this “all-to-all” connection, while AMD uses Infinity Fabric. This internal speed is why we treat the node as the basic building block of any cluster.

  • NVLink is the point-to-point link between GPUs, running at 900 GB/s on Hopper and 1,800 GB/s on Blackwell.
  • NVSwitch sits on top of NVLink and provides all-to-all communication, so every GPU in the node can talk to every other at full rate.
  • AMD Instinct uses Infinity Fabric for the equivalent job.

This is why the node is the natural unit of an HPC architecture, and it explains the oldest distinction in the field.

ModelWhat it meansWhere you see it
Shared memoryEvery processor addresses one common pool of memory. Simple to program, limited by how much you can physically attachInside a node. Programmed with OpenMP
Distributed memoryEach node has its own memory. Nodes exchange data by sending messages explicitlyAcross nodes. Programmed with MPI
HybridShared memory within the node, distributed memory across nodesEvery real cluster today. OpenMP or CUDA inside, MPI or NCCL across

Almost every modern system uses a “hybrid” model. This means you write code that treats the GPUs inside a node as a single, tightly-knit unit, while being very careful about how you send data over the network between different nodes.

The word “expensive” is doing real work in that sentence.

Layer 3. The Fabric: Connecting the Dots

When you move from communicating inside a node to communicating between nodes, speed takes a significant hit. Bridging this gap is the real challenge of HPC architecture.

Inside a Blackwell node, GPUs talk at 1,800 GB/s. The best node-to-node link available drops that to roughly 100 GB/s. That is an order of magnitude, and designing around it is most of what makes HPC architecture a discipline rather than a shopping list.

Why speed matters: In AI training, nodes constantly need to pause and “agree” on the data they’ve processed (a process called an all-reduce). If the network is slow, your expensive GPUs will sit idle waiting for updates. On a high-speed InfiniBand network, this might take 3 seconds; on standard Ethernet, it could take 20 seconds or more. This idle time is a common reason for poor cluster performance.

A collective runs at the speed of its slowest path.

Work the arithmetic. A 70-billion-parameter model in bf16 holds 140 GB of gradients. A ring all-reduce moves roughly twice that across each node’s uplink, so about 280 GB per node per step.

  • On an 800G InfiniBand fabric, roughly 100 GB/s, that exchange takes under 3 seconds.
  • On standard 100 gigabit Ethernet, roughly 12.5 GB/s, it takes over 20 seconds.

Frameworks overlap communication with computation, so you never pay the full number. But you cannot overlap bandwidth you do not have. On the wrong fabric the GPUs sit idle, and your dashboard reports it as low GPU utilisation rather than as a network problem. That is a common misdiagnosis in multi-node training.

The interconnect options

InterconnectWhat it isStrengthsWatch for
NVIDIA Quantum InfiniBandPurpose-built HPC fabric, now at 800GCongestion control managed in the fabric, predictable at scale, assumed by NVIDIA reference architecturesSingle supplier, higher cost per port
RoCE v2 over EthernetRDMA over Converged Ethernet, at 400G and 800GStandard switches, multi-vendor, no single supplier, works well with AMD InstinctNeeds careful DCQCN and PFC tuning to match InfiniBand behaviour
Omni-PathOriginally Intel, now Cornelis NetworksAn alternative HPC fabric with an installed base in traditional HPCSmaller ecosystem in AI clusters
Standard Ethernet, no RDMAOrdinary TCP/IP networkingCheap, universalNot viable for tightly coupled multi-node training

Whether you choose InfiniBand or RoCE v2 Ethernet depends on your needs. InfiniBand is the classic, high-performance choice often validated by NVIDIA, while open Ethernet offers more flexibility and vendor independence. The most important thing is that your provider gives you a choice that fits your workload.

Topology, and the ratio nobody volunteers

Switches are arranged in a topology. Two names dominate.

  • Fat-tree, also called leaf-spine. Nodes connect to leaf switches, leaf switches connect upward to spine switches. Bandwidth increases as you move up the tree, so the upper levels can carry the traffic of everything beneath them. This is what almost every AI cluster uses.
  • Dragonfly. Switches are grouped, groups are richly interconnected, and most traffic stays local. It reduces cable count and cost at very large scale, and appears more in national laboratory systems than in commercial AI clusters.

Whichever topology you get, one number decides what you actually receive.

  • Non-blocking, or 1:1. The uplink capacity leaving a leaf switch matches the capacity of the ports beneath it. Every node can transmit at full rate simultaneously.
  • Oversubscribed, often 3:1. The uplink carries a third of the capacity beneath it. This is a sensible, cost-saving design for web servers, which rarely all transmit at once. In an all-reduce every node transmits at once, so you receive a third of the fabric you thought you bought.

A “non-blocking” (1:1) network ensures every node can talk at full speed at the same time, which is exactly what you want for HPC.

Designing the network fabric for your HPC cluster? Talk to an AI Expert about Neysa’s non-blocking InfiniBand and RoCE v2 cluster architectures.

Layer 4. Storage: Fuel for Your GPUs

To keep thousands of processes running at once, you need a parallel file system. This allows many servers to read and write data simultaneously without getting in each other’s way.

SystemTypeTypically used for
LustreOpen source, the traditional HPC standardNational labs, academic clusters, and managed as a service by most large clouds
IBM Storage Scale (GPFS)Commercial, POSIX parallel file systemEnterprise HPC, mixed workloads
WEKACommercial, NVMe-firstAI clusters where checkpoint and small-file performance matter
VAST DataCommercial, disaggregated shared-everythingLarge AI estates with mixed hot and warm data
DDNCommercial, including EXAScalerLarge training clusters, long-established in HPC
Object storageS3-compatibleDatasets and archives. Not a substitute for a parallel file system at cluster scale

Fast storage is vital because of “checkpointing.” Every hour or so, a cluster saves its entire state so it doesn’t lose work if something fails. A single checkpoint can be 1 TB or more. If your storage is slow, your entire cluster stays idle while it saves. Over a week-long run, the difference between slow and fast storage can save you hours of expensive compute time.

A 70-billion-parameter model does not checkpoint 140 GB. It checkpoints the full training state: bf16 weights, an fp32 master copy, and two Adam optimizer moments. That comes to roughly 12 to 14 bytes per parameter, so about 1 TB per checkpoint. Checkpoint hourly, which is normal at that scale, and a seven-day run writes it 168 times, with the entire cluster idle for each write.

That is the same GPUs, the same job, and a difference of four hours of paid cluster time. Storage throughput is a compute cost line, not an IT cost line.

If you want to dive deeper into how different types of storage compare, our guide on object vs block vs file storage is a great place to start.

Layer 5. The Software Stack: Making It All Work

Hardware is just expensive decoration without the right software. The “stack” includes the languages, libraries, and frameworks that allow your code to actually talk to the chips and the network.

LayerOptionsWhat it does
Parallel programming modelMPI, OpenMP, CUDA, ROCm, SYCLMPI passes messages across nodes. OpenMP parallelises within a node. CUDA and ROCm target NVIDIA and AMD GPUs
Collective librariesNCCL, RCCL, MPI collectivesImplement all-reduce, all-gather and broadcast efficiently over the fabric. This is where topology awareness pays off
FrameworksPyTorch, JAX, DeepSpeed, Megatron-LM, vLLM, SGLangWhat you actually write against. Most call NCCL underneath without you seeing it
ContainersDocker, Apptainer (formerly Singularity), Enroot with PyxisPackage the environment so a job runs the same on every node. Apptainer and Enroot are preferred on HPC because they run unprivileged under a batch scheduler
LanguagesC, C++, Fortran, PythonFortran and C still dominate legacy scientific codes. Python dominates AI, with the heavy lifting in compiled kernels underneath

Many clusters that suddenly stop working have run into a version mismatch.

Layer 6. The Scheduler: Directing Traffic

Slurm is the standard tool used to manage HPC jobs. You tell it what you need, and it finds the right physical nodes for your job as they become available. For more dynamic workloads, many teams are also using Kubernetes.

By default, Slurm does not know how those nodes are wired. You tell it, through a file called topology.conf, which describes which nodes sit beneath which switch.

It is vital that your scheduler “understands” your network layout. If it places parts of a job on nodes that are physically far apart in the network, your communication will slow down significantly without any obvious error message in the logs.

Three ways a cluster is handed to you:

  • Pure bare metal. Root access and your own kernel. Take it if you have a platform team and strong opinions.
  • Bare metal with Slurm. Batch queueing, fair-share policy and checkpoint recovery, configured before handover. The default for classical training runs.
  • Bare metal with Kubernetes. Container orchestration, better suited to workloads that mix training with serving or that scale dynamically. See Kubernetes worker nodes explained.

Slurm and Kubernetes are not rivals so much as different assumptions. Slurm assumes a queue of finite jobs competing for a fixed pool. Kubernetes assumes long-running services that scale with demand. Large sites increasingly run both.

Layer 7. Operations: The Secret to Finishing

This layer doesn’t usually show up in technical diagrams, but it’s often the most important. Computers at this scale don’t always fail cleanly—sometimes a chip just slows down slightly, which can drag down the performance of your entire project.

GPUs rarely fail cleanly. They throttle under thermal pressure, log correctable ECC errors, or slow just enough to lag their rank. In a synchronous training job, one slow rank sets the pace for the entire cluster.

Then there is outright node loss, and the arithmetic is unforgiving.

A single node might be reliable. Five hundred of them running for a week are not. Two numbers cap your exposure:

  • How long a failed node stays in your allocation. A four-hour replacement instead of thirty minutes loses three and a half hours of the whole cluster, not of one node.
  • What share of incidents get caught before you report them. The difference between a swap that happens while you sleep and one that starts when you file a ticket.

The goal of good operations is goodput: the actual amount of useful work you keep. This involves proactive monitoring and quick hardware replacement. You can learn more about this in our article on GPU monitoring.

Putting It All Together

These layers are all connected. For example, if your storage is slow, you might decide to save checkpoints less often. But if a node fails, you lose more work between those saves. You end up spending more time redoing calculations, even if your GPU graphs look “busy.”

  1. Your file system delivers 10 GB/s instead of 100, so each checkpoint takes 100 seconds instead of 10.
  2. To reduce the number of checkpoints, you lengthen the interval from one hour to four.
  3. A node fails on hour three of an interval, as it eventually will on a large cluster.
  4. You restart from a checkpoint that is three hours old, on top of the node replacement time.
  5. Your GPU utilisation graph looks fine throughout, because the GPUs really were busy. They were just busy redoing work.

Nothing in that sequence is a GPU problem, and no GPU benchmark would have found it.

Benchmarking: how you know the architecture works

BenchmarkWhat it measuresWhen to use it
HPL / LINPACKDense linear algebra across the whole machine. The basis of the TOP500 listProving a cluster works end to end. A TOP500 entry means the fabric has been exercised, not just described
HPCGSparse workloads, a harder and more realistic memory-bound testWhere LINPACK flatters the system
NCCL testsAll-reduce and all-gather bandwidth at your actual node countThe single most useful acceptance test for an AI cluster
MLPerf Training and InferenceStandardised model workloads, auditedComparing platforms on work that resembles yours
IO500Storage performance under HPC access patternsValidating layer 3 before you sign

Before you commit to a cluster, always run real-world tests at full scale. A network that works for 8 nodes might struggle at 64. Testing early ensures you get the performance you’re paying for.

For monitoring in production, Prometheus with Grafana and NVIDIA DCGM Exporter has largely replaced the older Ganglia and Nagios pairing, though both still run at established sites.

Build or Rent?

 Build your ownHPC as a service
CapitalLarge upfront, depreciating over 3 to 5 yearsOperating expense, matched to the project
Time to first job9 to 18 months including power, space and commissioningWeeks
Silicon refreshYou own the obsolescenceThe provider does
Utilisation riskYou pay for the troughCommitted base plus on-demand peaks
Expertise neededFacilities, networking, storage, Slurm, 24×7 operationsYour workload, and a good contract
Best whenContinuous high utilisation, extreme data-gravity or security constraintsAlmost everything else

Should you build your own or use a service? If you can keep a cluster busy more than 70% of the time for three years, building might be for you. Otherwise, HPC as a Service or using an AI neocloud is often much more cost-effective.

How Neysa Can Help

Neysa provides dedicated HPC clusters based in India. We are proud to be the only India-headquartered provider highly rated in the SemiAnalysis ClusterMAX evaluation, which tests real-world performance on rented clusters.

Mapped to the six layers above, every one is a choice rather than a default:

  • Compute. NVIDIA and AMD, current generation. Around 2,000 GPUs live today, with 8,192 air-cooled AMD Instinct MI350X landing as a single-site cluster and 9,216 air-cooled NVIDIA B300 staged between December 2026 and March 2027, on a path past 25,000 GPUs by 2027.
  • Fabric. NVIDIA Quantum-X800 InfiniBand at 800G where you want the validated reference architecture, or non-blocking open-standards RoCE v2 Ethernet at 400G and 800G where you want vendor independence. Both offered non-blocking.
  • Storage. WEKA, VAST or DDN, open-source software-defined storage, or your own licences running on Neysa hardware.
  • Software and scheduler. Pure bare metal, bare metal with Slurm pre-configured for batch and checkpoint recovery, or bare metal with Kubernetes.
  • Operations. A 99.9% uptime SLA held 100% for six months running, 30-minute first response on Severity-1 against a four-hour resolution target, a GPU-replacement SLA written into the contract with instances restored in 30 to 60 minutes, and roughly one incident in three caught by monitoring before the customer reports it.
  • Security and Compliance. We maintain the highest standards, including ISO and SOC 2 certifications. You can find more details on our infrastructure security page.

Our clients, like ITQ Technologies, TIFIN, and Innoviti, have seen significant cost savings and performance improvements by using our tailored HPC solutions.

Interested in seeing how we compare to others? Take a look at our list of the Top 10 HPC Cloud Providers in India.

See predictable GPU cloud pricing for your workload. Talk to an AI Expert about deploying on Neysa Velocis.


  • Stop Thinking in Models. Start Thinking in Stacks. 

    Infrastructure

    9 mins.

    Stop Thinking in Models. Start Thinking in Stacks. 

    The AI stack comprises multiple layers: infrastructure, data handling, models, orchestration, inference, and governance. Efficient integration of these layers is crucial for performance, cost management, and reliability in real-world applications.


  • The Infrastructure Gap Stalling BFSI

    Infrastructure

    5 mins.

    The Infrastructure Gap Stalling BFSI

    Fraud detection at most banks still happens after the transaction completes. Not because the models are slow. It’s because the infrastructure running those models can’t respond fast enough to catch fraud while the money is still moving.


  • Neysa Velocis: Solving The Compute Trilemma

    Infrastructure

    7 mins.

    Neysa Velocis: Solving The Compute Trilemma

    There’s no single button that flips all three to “best”. Is there a pragmatic approach to treat the trilemma as a planning tool? This blog uncovers the approach for you.

SHARE