credit.slurm#

SLURM analog of credit.pbs.

Generates and (optionally) submits SLURM batch scripts for CREDIT training and rollout jobs. The public API mirrors credit.pbs so callers can swap from credit.pbs import ... for from credit.slurm import ... with no other changes.

All launchers use torchrun rather than mpirun/mpiexec: on SLURM the LOCAL_RANK / RANK / WORLD_SIZE variables that credit.distributed.get_rank_info() reads are set by torchrun, whereas srun alone only exports SLURM_PROCID (which is not recognized). Single node jobs run torchrun directly on the batch node; multi-node jobs use srun to launch one torchrun per node with a c10d rendezvous on the head node.

Config is read from a slurm: section if present, otherwise the pbs: section is reused (SLURM keys map as project/account -> –account, queue/partition -> –partition, ngpus -> GPUs per node, ncpus -> –cpus-per-task, mem -> –mem, walltime -> –time). Optional keys: gpu_type (adds --gres=gpu:<type>:<n>), constraint (--constraint), qos (--qos), modules (str or list, passed to module load), and env_setup (str or list of extra shell lines).

GPU request style differs by site: generic SLURM clusters request GPUs with --gres=gpu:N, but Perlmutter (NERSC) rejects that (“Job request does not match any supported policy”) and instead selects GPU nodes via --constraint=gpu + --qos + --gpus-per-node, needs no --partition or --mem line, and requires a _g account suffix. Perlmutter is detected from NERSC_HOST, a constraint config key, or cluster: perlmutter, and the correct directives are emitted automatically.

Attributes#

Functions#

launch_script(config_file, script_path[, launch, backend])

Generate and optionally submit a single-node SLURM script using torchrun.

launch_script_mpi(config_file, script_path[, launch, ...])

Generate and optionally submit a multi-node SLURM script using torchrun.

launch_script_torchrun(config_file, script_path[, ...])

Generate and optionally submit a SLURM script using torchrun.

get_num_cpus()

Return the number of CPUs available to the current job.

Module Contents#

credit.slurm.logger#
credit.slurm.launch_script(config_file, script_path, launch=True, backend='nccl')#

Generate and optionally submit a single-node SLURM script using torchrun.

Parameters:
  • config_file (str) – Path to the YAML configuration file.

  • script_path (str) – Path to the script that will be executed by the SLURM job.

  • launch (bool, optional) – If True, submit the job with sbatch. Defaults to True.

  • backend (str, optional) – Backend for distributed training. Defaults to ‘nccl’.

credit.slurm.launch_script_mpi(config_file, script_path, launch=True, backend='nccl')#

Generate and optionally submit a multi-node SLURM script using torchrun.

Uses srun to launch one torchrun per node with a c10d rendezvous on the head node – the SLURM equivalent of the mpiexec launcher in credit.pbs.

Parameters:
  • config_file (str) – Path to the YAML configuration file.

  • script_path (str) – Path to the script that will be executed.

  • launch (bool, optional) – If True, submit the job with sbatch. Defaults to True.

  • backend (str, optional) – Backend for distributed training. Defaults to ‘nccl’.

credit.slurm.launch_script_torchrun(config_file, script_path, launch=True, backend='nccl')#

Generate and optionally submit a SLURM script using torchrun.

Preferred for FSDP2 / v2-parallelism jobs. Single-node jobs use c10d + localhost; multi-node jobs broadcast the head node IP for the rendezvous endpoint via srun.

Parameters:
  • config_file (str) – Path to the YAML config file.

  • script_path (str) – Path to the training script (e.g., applications/train_gen2.py).

  • launch (bool) – If True, submit with sbatch. Defaults to True.

  • backend (str) – torch.distributed backend. Defaults to ‘nccl’.

credit.slurm.get_num_cpus()#

Return the number of CPUs available to the current job.

Inside a SLURM allocation this reads SLURM_CPUS_ON_NODE; otherwise it falls back to os.cpu_count().

credit.slurm.config_file = '../config/vit2d.yml'#