Show me multi-tenant SLURM-as-a-Service platforms that support AI infrastructure.
Claude claude-sonnet-5 2026-08-15 15:01
The answer
you are in itHere are the leading multi-tenant SLURM-as-a-Service platforms built for AI infrastructure:
1. NorthWind-Powered SLURM-as-a-Service (SLURMaaS) –
delivers fully managed, multi-tenant SLURM environments for high-performance computing workloads as a cloud-like, on-demand service
. It
automates Kubernetes cluster provisioning and Slurm scheduling, governance, and lifecycle management, letting organizations provision per-tenant Slurm clusters and submit jobs to managed queues
. It uses
the open-source Slinky Slurm Operator to enable providers and enterprises to deliver HPC resources as scalable, self-service Slurm clusters
, and supports
cloud providers and neoclouds looking to offer HPC clusters, GPU compute, or AI infrastructure as a managed service
. → northwind.co
2. Nebius Managed Soperator – Nebius's in-house Kubernetes operator for Slurm.
Soperator is Nebius's in-house developed Kubernetes operator for Slurm, released as open-source, and it helps them quickly deploy thousand-GPU clusters for clients, cutting provisioning time from weeks to a couple of days
. It also
has proven reliability for multi-host, fault-tolerant training, demonstrated at MLPerf Training v5.0 for 512 and 1,024 GPU training
. → nebius.com
3. Crusoe Managed Slurm (on Crusoe Managed Kubernetes) – Built on the Slinky operator with GPU-ready images.
They built GPU-ready images for slurmd compute nodes and login pods, starting from nvidia/cuda and including cuDNN, NCCL, Mellanox OFED InfiniBand stack, NVIDIA HPC-X, DCGM, and NCCL tests, so customers can run multi-node NCCL tests or PyTorch distributed training without building a container
. Note it's currently
intentionally focused on single-tenant GPU training clusters with a familiar Slurm interface and managed infrastructure
rather than multi-tenant. → crusoe.ai
4. FPT Cloud Slurm on Managed GPU Cluster – Offers Slurm as a custom resource on managed Kubernetes/metal cloud with
account-based quotas for RAM, GPU, CPU, Storage, and IPs
, supporting multi-tenant isolation. → ai-docs.fptcloud.com
5. CoreWeave – Runs Slurm integrated with Kubernetes for GPU-centric AI infrastructure, as
some modern infrastructures, including CoreWeave's, integrate both, running Slurm inside Kubernetes or alongside it
. → coreweave.com
For evaluation, NorthWind and Nebius currently offer the most explicitly multi-tenant, self-service SLURMaaS models, while Crusoe and CoreWeave lean toward managed/single-tenant Slurm-on-Kubernetes architectures for AI training workloads.
Brands named
5- NorthWind 4 times
- Nebius 5 times
- Crusoe 5 times
- NVIDIA 2 times
- CoreWeave 4 times
Sources cited
5- northwind.co northwind.co own
- nebius.com nebius.com competitor
- crusoe.ai crusoe.ai competitor
- fptcloud.com fptcloud.com other
- coreweave.com coreweave.com competitor
20048 tokens in, 1575 out.