Separate pools mean duplicated headroom.
Training clusters, inference endpoints, and agent evaluation each carry their own peak headroom. When they cannot share capacity, the physical fleet must size for three separate peaks instead of one.
One pool, governed by your priorities.
Nucleaton places training and inference into a single resource pool with policies you define. Guaranteed capacity is protected. Eligible spare capacity absorbs inference demand. When high-priority training arrives, supplemental serving yields.
Inference is not ordinary batch backfill.
Live inference demand arrives as individual requests, not queued batch jobs. Nucleaton's inference endpoints scale based on real request rate and queue depth. When capacity must return to training, the system can reclaim immediately or stop new routing and drain active requests before releasing GPUs.
Better economics from shared capacity.
Use committed training capacity before buying incremental inference capacity.
Reduce duplicated headroom across workload silos.
Absorb temporary inference peaks without dedicated serving clusters.
Run agents and evaluation on eligible spare capacity.
No universal savings percentage. Actual opportunity depends on workload timing, topology, GPU compatibility, model size, and priority constraints.
What determines the opportunity?
The amount of capacity available for sharing depends on your specific workload mix:
- Workload timing and diurnal patterns
- Cluster topology and GPU compatibility
- Model size and parallelism requirements
- Priority constraints and service guarantees
- Inference demand profile
Existing Slurm environments can be integrated through assisted deployment with appropriate administrative access.