
Looking for a reliable and fault-tolerant Slurm Workload Manager setup for your High-Performance Computing (HPC) environment?
I provide a professional service to install and configure Slurm in High Availability (HA) mode across multiple nodes (2 or more) to ensure job scheduling continues running smoothly even in case of a node failure.
🔧 What you will get:
✔️ Installation of Slurm controller (Primary + Backup) and compute nodes (CentOS/Rocky/Ubuntu)
✔️ HA configuration of Slurm controller using tools like Corosync + Pacemaker (or native Slurm HA if supported)
✔️ Setup of Munge authentication service
✔️ Configuration of compute nodes (slurmd)
✔️ Creation of partitions and job queue ready for production usage
✔️ Sample job submission examples to verify proper working
✔️ Ensuring Slurm services auto-start on system reboot
✔️ Shared storage configuration guidance (e.g., NFS, Lustre) if required
✔️ Failover configuration for slurmctld to the secondary node
✔️ Detailed instruction document
✔️ Basic post-install troubleshooting
🎯 Ideal for:
💡 Why choose me?
With deep expertise in HPC cluster management, Slurm, HA architecture, xCAT, and Linux systems, I ensure a highly available and reliable Slurm job scheduling solution tailored to your environment.
⚡ Delivery:
✔️ Setup in 1 day, depending on complexity
✔️ Post-installation support for 1 week (via chat)
👉 Message me now and ensure your HPC jobs never stop with a fully HA Slurm setup!