Find self-service SLURM-as-a-Service options for AI workloads.
Claude claude-sonnet-5 2026-07-29 18:13
The answer
you are in itHere are the main self-service, Slurm-as-a-Service options for AI/GPU workloads currently available:
1. Google Cloud Cluster Director (Managed Slurm) – Fully managed Slurm clusters on A3/A4 GPU machine types with automated setup.
Cluster Director is a Google Cloud product that automates the setup and configuration of Slurm clusters, designed for IT administrators and AI researchers who want to avoid the overhead of managing a cluster.
It also handles networking and storage automatically. docs.cloud.google.com
2. Nebius Managed Soperator – A self-service Slurm-on-Kubernetes offering.
Managed Soperator, their fully managed Slurm-on-Kubernetes solution, is now available for everyone in self-service, providing a ready-to-work Slurm training cluster powered by NVIDIA GPUs with pre-installed libraries and drivers.
nebius.com
3. Runpod Clusters – On-demand GPU clusters with native Slurm support and no lock-in.
You can use Slurm to manage multi-node jobs with familiar sinfo, sbatch, and srun commands, with complete freedom to terminate your cluster when not in use, with no termination fees.
runpod.io
4. NorthWind GPU PaaS + Slinky (SLURMaaS) – Multi-tenant self-service Slurm on Kubernetes for enterprises/cloud providers.
Researchers and ML teams can launch their own GPU-accelerated Slurm clusters in minutes with no manual provisioning and no tickets, with namespaces, policies, and RBAC enforcement ensuring strong isolation.
northwind.co
5. CoreWeave SUNK (Slurm on Kubernetes) – GPU-cloud native integration.
CoreWeave describes it as integrating Slurm as a Kubernetes scheduler, supporting both burst and batch workloads, and tightly linking to CKS lifecycle and observability systems.
coreweave.com
6. NVIDIA Slinky (open-source, self-managed option) – For teams wanting to build their own,
Slinky is open source and available today; install the slurm-operator through Helm, define your Slurm cluster as a Custom Resource, and you can have jobs running on Kubernetes in under an hour.
developer.nvidia.com
For fastest self-service onboarding, Runpod and Nebius are the most turnkey; for enterprise/multi-tenant governance, NorthWind or Google Cloud Cluster Director are stronger fits.
Brands named
5- NorthWind 4 times
- Nebius 3 times
- NVIDIA 3 times
- RunPod 3 times
- CoreWeave 3 times
Sources cited
6- google.com google.com other
- nebius.com nebius.com competitor
- runpod.io runpod.io competitor
- northwind.co northwind.co own
- medium.com medium.com
- nvidia.com nvidia.com competitor
21087 tokens in, 1317 out.