Slurm as the foundation, Nucleaton as the platform.
Slurm remains the authoritative resource scheduler for GPU allocation. Nucleaton operates on top, automating cluster deployment, enforcing resource policy, managing training jobs, and placing inference workloads into eligible spare capacity.
An optional Kubernetes overlay provides inference endpoints, agent hosting, and model serving. Both Slurm and K8s draw from the same physical fleet, coordinated by Nucleaton's resource policy layer.
Purpose-built for GPU training at scale.
Slurm is the de facto workload manager for large-scale GPU training. It understands job arrays, node partitions, hardware topology, and fair scheduling across thousands of GPUs. Nucleaton doesn't replace Slurm; it extends it with automated deployment, resource policy, converged inference placement, and customer-facing service delivery.
Each layer adds capability. None replaces your scheduler.
Inference endpoints, agents, model hosting, and branded tenant portals.
Guaranteed, flexible, and reclaimable capacity tiers governed by operator-defined rules.
Priority-aware scheduling that places training and inference into one physical pool.
OS, drivers, networking, parallel filesystems, and Slurm configuration, deployed in minutes.
From tenant contract to physical fleet.
For datacenter operators, Nucleaton extends northward to meet customer contracts and southward to manage physical infrastructure. The same policy engine that places workloads also connects to Redfish, BMC interfaces, and datacenter management systems.
This creates a single control loop: customer commitments become enforceable resource policies, resource policies drive placement decisions, and placement decisions respond to real hardware state.
Northbound commitments, southbound control.
The northbound interface handles tenant onboarding, service level definitions, and capacity commitments. The southbound interface manages compute nodes, GPU health, scheduling partitions, and physical infrastructure state.
Datacenter integrations are currently in beta.
Supported management systems are expanding.
Existing Slurm clusters can be integrated through assisted deployment with appropriate administrative access.