credit.slurm#
SLURM analog of credit.pbs.
Generates and (optionally) submits SLURM batch scripts for CREDIT training and
rollout jobs. The public API mirrors credit.pbs so callers can swap
from credit.pbs import ... for from credit.slurm import ... with no other
changes.
All launchers use torchrun rather than mpirun/mpiexec: on SLURM the
LOCAL_RANK / RANK / WORLD_SIZE variables that
credit.distributed.get_rank_info() reads are set by torchrun, whereas
srun alone only exports SLURM_PROCID (which is not recognized). Single
node jobs run torchrun directly on the batch node; multi-node jobs use srun
to launch one torchrun per node with a c10d rendezvous on the head node.
Config is read from a slurm: section if present, otherwise the pbs:
section is reused (SLURM keys map as project/account -> –account,
queue/partition -> –partition, ngpus -> GPUs per node, ncpus ->
–cpus-per-task, mem -> –mem, walltime -> –time). Optional keys:
gpu_type (adds --gres=gpu:<type>:<n>), constraint (--constraint),
qos (--qos), modules (str or list, passed to module load), and
env_setup (str or list of extra shell lines).
GPU request style differs by site: generic SLURM clusters request GPUs with
--gres=gpu:N, but Perlmutter (NERSC) rejects that (“Job request does not
match any supported policy”) and instead selects GPU nodes via
--constraint=gpu + --qos + --gpus-per-node, needs no --partition
or --mem line, and requires a _g account suffix. Perlmutter is detected
from NERSC_HOST, a constraint config key, or cluster: perlmutter, and
the correct directives are emitted automatically.
Attributes#
Functions#
|
Generate and optionally submit a single-node SLURM script using torchrun. |
|
Generate and optionally submit a multi-node SLURM script using torchrun. |
|
Generate and optionally submit a SLURM script using torchrun. |
Return the number of CPUs available to the current job. |
Module Contents#
- credit.slurm.logger#
- credit.slurm.launch_script(config_file, script_path, launch=True, backend='nccl')#
Generate and optionally submit a single-node SLURM script using torchrun.
- Parameters:
config_file (str) – Path to the YAML configuration file.
script_path (str) – Path to the script that will be executed by the SLURM job.
launch (bool, optional) – If True, submit the job with
sbatch. Defaults to True.backend (str, optional) – Backend for distributed training. Defaults to ‘nccl’.
- credit.slurm.launch_script_mpi(config_file, script_path, launch=True, backend='nccl')#
Generate and optionally submit a multi-node SLURM script using torchrun.
Uses
srunto launch one torchrun per node with a c10d rendezvous on the head node – the SLURM equivalent of the mpiexec launcher incredit.pbs.- Parameters:
config_file (str) – Path to the YAML configuration file.
script_path (str) – Path to the script that will be executed.
launch (bool, optional) – If True, submit the job with
sbatch. Defaults to True.backend (str, optional) – Backend for distributed training. Defaults to ‘nccl’.
- credit.slurm.launch_script_torchrun(config_file, script_path, launch=True, backend='nccl')#
Generate and optionally submit a SLURM script using torchrun.
Preferred for FSDP2 / v2-parallelism jobs. Single-node jobs use c10d + localhost; multi-node jobs broadcast the head node IP for the rendezvous endpoint via
srun.- Parameters:
config_file (str) – Path to the YAML config file.
script_path (str) – Path to the training script (e.g., applications/train_gen2.py).
launch (bool) – If True, submit with
sbatch. Defaults to True.backend (str) – torch.distributed backend. Defaults to ‘nccl’.
- credit.slurm.get_num_cpus()#
Return the number of CPUs available to the current job.
Inside a SLURM allocation this reads
SLURM_CPUS_ON_NODE; otherwise it falls back toos.cpu_count().
- credit.slurm.config_file = '../config/vit2d.yml'#