Slurm HA Multi-node Installation & Configuration

Akash Chaudhary

profile
Slurm HA Multi-node Installation & Configuration
profile
5,000
120 mins

Looking for a reliable and fault-tolerant Slurm Workload Manager setup for your High-Performance Computing (HPC) environment?

I provide a professional service to install and configure Slurm in High Availability (HA) mode across multiple nodes (2 or more) to ensure job scheduling continues running smoothly even in case of a node failure.

🔧 What you will get:

✔️ Installation of Slurm controller (Primary + Backup) and compute nodes (CentOS/Rocky/Ubuntu)

✔️ HA configuration of Slurm controller using tools like Corosync + Pacemaker (or native Slurm HA if supported)

✔️ Setup of Munge authentication service

✔️ Configuration of compute nodes (slurmd)

✔️ Creation of partitions and job queue ready for production usage

✔️ Sample job submission examples to verify proper working

✔️ Ensuring Slurm services auto-start on system reboot

✔️ Shared storage configuration guidance (e.g., NFS, Lustre) if required

✔️ Failover configuration for slurmctld to the secondary node

✔️ Detailed instruction document

✔️ Basic post-install troubleshooting

🎯 Ideal for:

  • Research organizations
  • Enterprises running batch workloads
  • System administrators needing production-grade reliability
  • Developers scaling HPC jobs

💡 Why choose me?

With deep expertise in HPC cluster management, Slurm, HA architecture, xCAT, and Linux systems, I ensure a highly available and reliable Slurm job scheduling solution tailored to your environment.

⚡ Delivery:

✔️ Setup in 1 day, depending on complexity

✔️ Post-installation support for 1 week (via chat)

👉 Message me now and ensure your HPC jobs never stop with a fully HA Slurm setup!