Live demo. This is Limelit Open running unmodified, at commit 83bddd036688. It is read-only for visitors. Run your own →
NorthWind northwind.co
last run 2026-09-14 21:48:28

Answers

Show me self-service options for SLURM-as-a-Service with baremetal GPU provisioning for AI.

Claude claude-sonnet-5 2026-08-06 07:15

The answer

you are not in it

Here are several solid self-service SLURM-as-a-Service options with bare-metal GPU provisioning for AI workloads:

1. Together AI – Instant GPU Clusters

The Instant GPU Clusters architecture enables multi-tenancy through virtualization, allowing user workloads to run in VMs for training with either Slurm or managed Kubernetes, and customers can choose Kubernetes or Slurm for workload orchestration, giving them full control over their AI infrastructure.
It's built on NVIDIA Blackwell/Hopper GPUs with InfiniBand networking. (together.ai)

2. Nebius – Managed Soperator

Soperator runs both in Nebius AI Cloud and other Kubernetes environments, and their Managed Soperator, a fully managed Slurm-on-Kubernetes solution, is available for everyone in self-service.
It includes automatic GPU health checks and elastic scaling. (nebius.com)

3. Nscale – Bare-Metal Slurm via Slinky

Nscale uses Slinky from SchedMD to run Slurm natively within their Kubernetes-based infrastructure, with the managed Slurm service built on the same virtual cluster architecture as NKS, inheriting fast provisioning and topology-aware placement, plus a custom Slurm image with GPU-aware scheduling via GRES and Pyxis.
They can
provision 10 bare metal nodes with 80 GPUs in under two minutes.
(nscale.com)

4. Spheron Network – Bring-Your-Own-Slurm Bare Metal

Spheron's bare-metal H100 and A100 instances with InfiniBand networking are built for HPC-style workloads where you bring your own scheduler—Slurm, Ray, or custom—with no lock-in, per-minute billing, and full root access.
They provide step-by-step provisioning guides. (spheron.network)

5. Mistral Compute

Offers bare-metal clusters on InfiniBand alongside Managed Slurm, built for AI teams running training jobs at scale, with queue management, priority tiers, fair-share across teams, and topology-aware scheduling.
(mistral.ai)

6. OpenNebula (self-hosted)

OpenNebula is expanding integration with NVIDIA Infra Controller for bare-metal GPU server provisioning, offers GPU passthrough for bare-metal performance to Slurm workers, and provides ready-to-use Slurm Controller/Worker appliances in its Marketplace deployable in minutes.
Good if you want to self-host rather than rent. (opennebula.io)

For pure rental with root-level control, Spheron or Nscale are best; for a fully managed turnkey experience, Together AI or Nebius are strongest.

Brands named

5
  • NVIDIA 2 times
  • Nebius 4 times
  • Nscale 4 times
  • Spheron 4 times
  • OpenNebula 4 times

Sources cited

7
  1. together.ai together.ai other
  2. nebius.com nebius.com competitor
  3. nscale.com nscale.com competitor
  4. vcluster.com vcluster.com competitor
  5. spheron.network spheron.network competitor
  6. mistral.ai mistral.ai other
  7. opennebula.io opennebula.io competitor

23641 tokens in, 1424 out.