GPU failures across large training fleets require automatic remediation because manual intervention does not scale.
2
Crusoe runs Managed Slurm on Kubernetes so researchers keep Slurm workflows while platform teams gain Kubernetes health management and observability.
3
AutoClusters can replace a failed GPU node and resume checkpointed training in under 15 minutes without user action.
Summary
Connor Guerrero, Young Jeong, and Nikhil Gupta explain why Crusoe combines Slurm with Kubernetes for large GPU training clusters. Slurm provides gang scheduling, topology awareness, familiar sbatch workflows, and other features that suit distributed training. It is less suited to dynamic resource sharing, automated node health management, and detailed failure diagnosis. Crusoe runs Managed Slurm on its managed Kubernetes platform, using Kubernetes to manage infrastructure while Slurm remains the interface for researchers. The talk walks through AutoClusters handling an XID 79 GPU failure. The system notifies the user, stops work on the affected node, drains and replaces it, requeues the job, and lets the application reload its checkpoint. In the demo, replacing the node takes about five minutes, while the full process takes under 15 minutes. The same shared pool can also move GPUs between training and inference workloads as demand changes.
Large GPU fleets make manual failure recovery unsustainable
Connor Guerrero says customers run very large training workloads across thousands of GPUs, so failures are inevitable. Handling each failure by hand does not scale. Crusoe's approach combines Slurm and Kubernetes to automate remediation while keeping the workflow familiar for machine learning engineers and the operations model manageable for platform teams. The goal is to deal with hardware errors without requiring an engineer to diagnose and repair a node during an overnight incident.
Slurm fits distributed training because it understands tightly coordinated jobs
Young Jeong describes Slurm as a system built for researchers and high-performance computing. Its model fits modern training workloads that need tight collective communication across ranks and specialized training networks. Gang scheduling lets the required workers start together, while topology awareness can place work with the network in mind. Prologue and epilogue hooks support cluster validation. Researchers can also keep using familiar scripts and sbatch commands.
Traditional Slurm needs extra systems for dynamic workloads and node health
Jeong says modern AI work includes training, post-training evaluation, and inference, with resources changing as workloads change. Traditional Slurm is relatively static in its partitioning and related functions. GPU failures add another operational task. Teams must detect bad nodes, drain them, requeue jobs, and decide how to automate those steps. Slurm can provide pieces of this process, but customers may need scripts and additional tools to connect them.
Kubernetes adds the infrastructure controls that Slurm lacks
Nikhil Gupta describes Kubernetes as a mature cloud platform with tools for observability, networking, and security. It also provides mechanisms for maintaining services, including health handling, load balancing, and autoscaling. Crusoe saw teams arrive from either direction: Kubernetes teams wanted training, while Slurm teams were unfamiliar with Kubernetes frameworks. Running separate stacks would create extra infrastructure, operational work, and cost.
Managed Slurm keeps researcher and platform workflows separate
Crusoe runs Managed Slurm on Managed Kubernetes through its Slurm operator and the Slinky project. A single command creates the Slurm cluster, including its users, partitions, configuration, and storage. Researchers connect over SSH and use the Slurm interface without needing to know that Kubernetes is underneath. Platform teams see Slurm as another service inside their existing infrastructure, with shared observability, runbooks, and on-call processes rather than a second standalone stack.
A shared GPU pool can shift capacity between training and inference
When Slurm and Kubernetes use the same underlying GPU pool, resources can move between workloads. A platform can allocate more hardware to inference when user demand rises, then return those GPUs to training when inference demand falls. Finished Slurm jobs do not leave GPUs trapped in an idle training cluster. Kubernetes can schedule inference pods on the same nodes because they remain part of the shared node pool.
AutoClusters turns an XID 79 failure into an automatic replacement flow
For an XID 79 error that makes a GPU unusable, AutoClusters first notifies the user, although no action is required. The Slurm operator marks the node down and cancels its jobs. The process receives SIGTERM and has up to two minutes to save checkpoints or flush logs. The job is requeued. AutoClusters checks whether replacement is allowed, cordons and drains the Kubernetes node, removes it from the pool, and adds a healthy spare node. The job then starts again and loads its checkpoint.
Checkpointing lets training continue after a failed GPU
In the demo, Guerrero starts a PyTorch training job with sbatch on a pool containing two A100 nodes. He triggers an XID 79 event, after which GPU utilization drops and node replacement begins. The infrastructure replacement takes roughly five minutes. Application startup, model loading, and checkpoint loading account for most of the remaining time, making the full recovery a little under 15 minutes. The job resumes without a human repairing the cluster, and GPU utilization returns to normal.
"The full process from detecting a critical hardware error to getting a healthy node back into the node pool takes roughly five minutes with AutoClusters."Connor Guerrero12:18
Who should watch
You run distributed training on Slurm and need a way to handle failed GPU nodes without waking an engineer.
Your platform team supports both Kubernetes services and Slurm workloads and wants one infrastructure stack.
You want to understand how checkpointing, job requeueing, and node replacement fit together during a GPU failure.