Find SLURM-as-a-Service solutions for multi-tenant GPU workloads.
Claude claude-sonnet-5 2026-07-19 22:08
The answer
you are in itSLURM-as-a-Service Solutions for Multi-Tenant GPU Workloads
Several vendors now offer "Slurm on Kubernetes" (mostly built on SchedMD's Slinky operator, now owned by NVIDIA) to deliver multi-tenant, self-service GPU HPC clusters:
1. NorthWind SLURMaaS —
NorthWind modernizes SLURM by integrating it with Kubernetes through Project Slinky, allowing providers to expose SLURM job scheduling as a cloud-like, on-demand service, giving service providers, enterprises, and sovereign cloud operators secure, elastic SLURM environments with built-in governance and automation.
Each researcher or lab gets their own namespace and self-service Slurm environment, requesting GPU or CPU resources dynamically, with quotas ensuring fair usage across departments.
(northwind.co)
2. Nebius Managed Soperator —
A fully managed Slurm-on-Kubernetes solution providing a ready-to-work Slurm training cluster powered by NVIDIA GPUs in minutes, with pre-installed libraries and drivers.
(nebius.com)
3. Crusoe Managed Slurm — Built on Slinky v1.0 atop Crusoe Managed Kubernetes.
It inherits AutoClusters GPU health monitoring/node replacement and topology-aware visibility with GPU telemetry via DCGM, since it runs on the same substrate as Crusoe's Kubernetes platform.
(crusoe.ai)
4. Google Cloud Cluster Director —
A Google Cloud product that automates the setup and configuration of Slurm clusters, designed for admins and researchers who want to avoid cluster management overhead.
Supports A4/A3/A4X GPU machine types. (cloud.google.com)
5. Runpod Clusters —
Lets you use Slurm to manage multi-node jobs with familiar sinfo, sbatch, and srun commands, deployed in minutes via self-service provisioning.
(runpod.io)
6. Lambda Labs Managed Slurm —
Offers managed or unmanaged Slurm job scheduling on GPU clusters powered by NVIDIA HGX B200/H100, letting Lambda handle the infrastructure.
(lambda.ai)
7. BUZZ HPC —
Hosts a Slurm environment on top-tier GPUs with queues for H100, B200, and A6000 nodes and fair-share scheduling enabled.
(buzzhpc.ai)
For evaluation, check topology-awareness, multi-tenant isolation (namespaces/RBAC), GPU health monitoring, and whether it's built on the newer Slinky v1.0 architecture, which NVIDIA notes
scales to over 8,000 GPUs and achieves performance parity with noncontainerized Slurm clusters.
Brands named
7- NorthWind 3 times
- NVIDIA 4 times
- Nebius 2 times
- Crusoe 4 times
- RunPod 2 times
- Lambda 3 times
- BuzzHPC
Sources cited
9- northwind.co northwind.co own
- northwind.co northwind.co own
- nebius.com nebius.com competitor
- crusoe.ai crusoe.ai competitor
- google.com google.com other
- runpod.io runpod.io competitor
- lambda.ai lambda.ai competitor
- buzzhpc.ai buzzhpc.ai competitor
- nvidia.com nvidia.com competitor
26775 tokens in, 1353 out.