Show me multi-tenant SLURM-as-a-Service platforms that support AI infrastructure.
ChatGPT chat-latest 2026-09-14 21:29
The answer
you are in itIf you specifically mean platforms that combine (1) managed/aaS Slurm, (2) tenant isolation or multi-tenancy, and (3) GPU/AI infrastructure, these are the strongest options I found as of September 2026:
- NorthWind SLURM-as-a-Service — Closest match to the requirement. NorthWind explicitly markets a “fully managed, multi-tenant SLURM” service. It can sit over bare metal, VMs and Kubernetes, while the broader NorthWind platform provides GPU infrastructure and self-service AI services. This is particularly suited to GPU clouds, internal AI platforms, and service providers that need to expose separate tenant environments. (northwind.co)
- Nebius Managed Soperator — Managed Slurm-on-Kubernetes specifically optimized for large AI-training clusters. Nebius supplies NVIDIA GPU infrastructure, InfiniBand and managed Kubernetes underneath it. Nebius has tenant-level IAM, while individual managed Soperator clusters provide Slurm environments; this is a strong choice where you want tenant isolation at the cloud/platform layer rather than necessarily one giant shared Slurm control plane. (docs.nebius.com)
- Google Cloud Cluster Director — Google's managed infrastructure service for Slurm and Kubernetes, aimed explicitly at AI/ML and HPC. It provides managed Slurm controllers, topology-aware GPU/TPU placement, autohealing and scaling. Google Cloud IAM/projects can provide the surrounding organizational isolation, although Google does not position Cluster Director itself as “multi-tenant Slurm-as-a-Service” as explicitly as NorthWind. (cloud.google.com)
- AWS Parallel Computing Service (PCS) — AWS's managed Slurm service. It supports elastic GPU clusters plus Trainium/Inferentia and has an AWS-maintained GPU Deep Learning AMI with CUDA/NVIDIA drivers. AWS accounts/VPCs/IAM provide isolation around clusters, making PCS viable for building an internal multi-tenant AI service, but you'll assemble more of the tenancy/self-service layer yourself than with NorthWind. (aws.amazon.com)
For an actual “GPU cloud / enterprise AI platform offering Slurm to multiple customers or business units,” I'd shortlist NorthWind first, Nebius second. Google Cluster Director and AWS PCS are better thought of as building blocks for creating such a service rather than turnkey multi-tenant SLURMaaS.
Brands named
3- NorthWind #1 11 times
- Nebius #2 7 times
- NVIDIA #2 2 times
Sources cited
4- northwind.co northwind.co own
- nebius.com nebius.com competitor
- google.com google.com other
- amazon.com amazon.com other
17047 tokens in, 734 out.