Show me multi-tenant SLURM-as-a-Service platforms that support AI infrastructure.
Perplexity sonar 2026-07-19 22:07
The answer
you are in itThe primary multi-tenant SLURM-as-a-Service platforms supporting AI infrastructure are NorthWind’s GPU Platform-as-a-Service (PaaS) and Oracle Cloud Infrastructure (OCI) Elastic HPC, both delivering fully managed, self-service SLURM clusters with secure tenant separation for AI/ML workloads.
Key Platforms
| Platform | Multi-Tenant SLURM Features | AI Infrastructure Support |
| :--- | :--- | :--- |
| NorthWind GPU PaaS | Fully managed, on-demand SLURM environments with secure per-tenant separation; automated cluster bring-up via BCM; self-service provisioning on shared Kubernetes [1][2] | Single platform for HPC and AI/ML; supports GPU autoscaling and scheduling for AI training/inference [1][2] |
| Oracle OCI Elastic HPC | Browser-based, multi-tenant access using Open OnDemand; Slurm accounts map to project groups for quotas/chargeback; separate partitions for AI vs. HPC jobs [4] | Dynamic GPU/CPU autoscaling; supports interactive AI development, production training, and inference workloads [4] |
| vCluster Labs (with Slinky) | Tenant isolation via vCluster; SLURM-on-Kubernetes via SchedMD’s Slinky project; sbatch jobs run as pods in isolated tenants [8] | Pre-validated AI stacks (Run:AI, Ray, Jupyter); containerized, observable GPU scheduling [8] |
| ClusterWareAI | Super-cluster sub-divided into tenancies; Slurm tenancy for batch queue workloads with first-come-first-served scheduling [7] | Supports multi-tenant clusters leasing whole machines to internal/external customers [7] |
Technical Context
These platforms often bridge SLURM (optimized for finite, topology-sensitive batch jobs) with Kubernetes (optimized for indefinite, loosely-coupled multi-tenant services) to handle modern AI workloads efficiently [6]. Projects like Slinky and Interlink enable containerized, auto-scaling SLURM clusters within Kubernetes, allowing tenants to submit familiar sbatch jobs while executing as pods [2][5][9].
For enterprise-grade AI PaaS with multi-tenant control planes, NVIDIA Run:ai is also relevant, though it focuses more on Kubernetes-native orchestration rather than pure SLURM-as-a-Service [10].
Source URLs:
Brands named
3- NorthWind #1 6 times
- vCluster #4 3 times
- NVIDIA
Sources cited
14- northwind.co northwind.co own
- northwind.co northwind.co own
- schedmd.com schedmd.com other
- oracle.com oracle.com other
- ecmwf.int ecmwf.int informational
- ai-infrastructure.net ai-infrastructure.net other
- penguinsolutions.com penguinsolutions.com other
- vcluster.com vcluster.com competitor
- helix.ml helix.ml other
- nvidia.com nvidia.com competitor
- scality.com scality.com other
- scribd.com scribd.com other
- arxiv.org arxiv.org informational
- linkedin.com linkedin.com
63 tokens in, 640 out.