Jobs
Overview
Jobs on your cluster are scheduled and run using the Slurm workload manager. Resources can be allocated through two main mechanisms:
Interactive (
salloc): allocate resources and run commands directly, useful for development and debugging.Batch (
sbatch): submit a script that runs when resources become available, the standard approach for production workloads.
Within an allocation you can then use the srun command to execute job steps and further distribute work across your allocation.
For command reference, see the official Slurm documentation.
Inspecting the Cluster
Before we submit jobs, it's useful to understand how we can inspect the state of the cluster and queue.
sinfo
sinfo is the basic mechanism for inspecting the state of the resources available
sinfosqueue
squeue is how we can inspect the state of the queue. You can use the -u flag to check the status of your current running jobs.
squeue -u $USERResource Specification
Though there are some differences in how resources are specified with the salloc, sbatch, and srun commands, the common directives are generally the same.
--nodes / -N
Number of nodes
--ntasks-per-node
Number of tasks (processes) per node
--gpus-per-node
GPUs per node
--cpus-per-task
CPU cores per task
--time
Wall-clock time limit (HH:MM:SS)
--nodelist
Run on specific nodes
This is not a comprehensive list of the available flags, please reference Slurm documentation for full man pages.
All jobs must request GPU resources explicitly. Reasonable defaults are established to subdivide CPU cores and memory based on the GPU allocation, but these can be overridden per job.
Job Steps with srun
Within the following allocation methods, it is useful to understand how job steps work and how they can be used to maximize allocations.
By default, srun commands will inherit the entire resource allocation for all subcommands. This is useful for submitting monolithic jobs, but can be tuned to instead subdivide resources within an allocation for multiple tasks. Consider the following examples in ways srun can be used from within an allocation:
Interactive Allocation with salloc
To launch an interactive allocation within Slurm, use the salloc command.
The salloc session holds the allocation open. When you exit the salloc shell, the allocation is released and the node becomes available to other jobs.
For an interactive shell directly on the worker pod, you can add the following srun:
To drop into an interactive shell inside a container, use apptainer shell with srun (don't forget the --pty flag):
This allocates a worker pod, pulls (or uses a cached) container image, and drops you into an interactive shell inside it with GPUs available. Use a local .sif file instead of a docker:// URI if you have already pulled the image. For more on building and running Apptainer images, including multi-node jobs and networking, see Containers and Modules.
Batch Allocation with sbatch
Batch jobs are submitted with sbatch and run when the scheduler grants the allocation. Resource requests, environment setup, and the actual workload are all defined in the job script. Any #SBATCH directive in the script can also be overridden at submission time by passing the corresponding flag directly to sbatch.
Basic structure
Submit with:
View output once the job runs:
Multi-node distributed training example
The pattern below works for PyTorch torch.distributed.run (torchrun) across multiple nodes. The first node allocated by Slurm acts as the rendezvous host.
Example scripts are available on the cluster under /opt/tw/examples/libexec/.
SSH
Though it's recommended that you use the above mechanisms for submitting work, you are able to access any node within your allocation via SSH for debugging purposes.
MPI (PMIx)
The cluster uses PMIx as the default MPI launch interface (MpiDefault=pmix). Open MPI and other PMIx-compatible MPI libraries work with srun without needing mpirun.
The following environment is set cluster-wide and applied automatically to all login and compute sessions:
Running an MPI job
Use srun directly. Slurm handles process launch and PMIx initialization:
We recommend using srun rather than mpirun or mpiexec, as it integrates directly with Slurm's process placement and PMIx initialization.
RCCL collective communication
For GPU collective benchmarks and validation, RCCL tests are pre-installed on compute nodes. A minimal all-reduce test across 2 nodes:
For containerized jobs (Pyxis/Enroot, Apptainer) and environment modules, see Containers and Modules.
Last updated

