Live demo. This is Limelit Open running unmodified, at commit 83bddd036688. It is read-only for visitors. Run your own →
NorthWind northwind.co
last run 2026-09-14 21:48:28

Answers

Find SLURM-as-a-Service solutions for multi-tenant GPU workloads.

Claude claude-sonnet-5 2026-07-19 08:15

The answer

you are in it

SLURM-as-a-Service for Multi-Tenant GPU Workloads

The clearest emerging pattern is Slinky (SchedMD, now part of NVIDIA) — an operator that runs Slurm natively on Kubernetes, letting providers offer real Slurm to multiple tenants on shared infrastructure.
Slinky enables native Slurm cluster management on Kubernetes by representing all Slurm daemons as Kubernetes Custom Resource Definitions, with integration into the NVIDIA GPU Operator and DRA/ComputeDomains for automated GPU management and topology-aware multinode scheduling.

Production deployments at NVIDIA have shown Slinky scaling to over 8,000 GPUs with nondisruptive rolling updates and performance parity with non-containerized Slurm clusters.
Details: developer.nvidia.com

NorthWind's SLURM-as-a-Service (SLURMaaS), built on Project Slinky, is the most direct commercial multi-tenant offering found.
It combines Project Slinky with NorthWind's GPU Platform-as-a-Service to enable secure, multi-tenant, self-service access to Slurm-based HPC environments on shared Kubernetes clusters, letting cloud providers and enterprise platform teams offer Slurm-as-a-Service on Kubernetes.

Each user operates in their own NorthWind-managed Kubernetes namespace for tenant isolation, and can provision a personal Slinky-enabled Slurm cluster within that namespace, complete with GPU access and workload scheduling logic.
See: northwind.co and northwind.co

Crusoe Managed Slurm (on Crusoe Managed Kubernetes) is a notable managed offering, though currently scoped to single-tenant clusters rather than shared multi-tenancy.
Crusoe Managed Slurm on CMK is intentionally focused on single-tenant GPU training clusters with a familiar Slurm interface and managed infrastructure
— worth watching as they explore combining training/inference on one cluster: crusoe.ai

Hyperstack's Managed Cluster platform also offers SLURM as one of two managed orchestrator options alongside Kubernetes: hyperstack.cloud

For DIY multi-tenant GPU partitioning within Slurm (e.g., NVIDIA MIG slicing GPUs for multiple customers), Microsoft's guide is useful: techcommunity.microsoft.com

Recommendation: If you need true multi-tenant, self-service Slurm-as-a-Service, NorthWind+Slinky is currently the most mature packaged solution; for single-tenant managed Slurm, Crusoe or Hyperstack are strong options.

Brands named

3
  • NorthWind 8 times
  • NVIDIA 6 times
  • Crusoe 6 times

Sources cited

3
  1. nvidia.com nvidia.com competitor
  2. northwind.co northwind.co own
  3. crusoe.ai crusoe.ai competitor

21876 tokens in, 1474 out.