HPC systems (Slurm, Slingshot)¶
Running DDStore under Slurm on Cray Slingshot systems (Frontier, Perlmutter):
the example job scripts, which --network options a layout needs, and what
to do when RDMA can’t connect.
Do you need --network flags?¶
Every srun step gets its own Slingshot VNI (network isolation ID), and two
endpoints can only reach each other on the same VNI. Whether you need
#SBATCH --network=single_node_vni,job_vni depends on how steps talk:
Layout |
Perlmutter, no flags |
Perlmutter, |
Frontier |
|---|---|---|---|
One |
works |
works |
works; |
Core and extra as separate one-node steps ( |
works: one-node steps get no VNI and share the default one |
needs |
needs both flags + the wrapper |
Core and extra on the same nodes, or steps spanning several nodes |
extra can’t reach core |
needs |
second step fails to launch |
The Perlmutter columns were measured on 2 nodes (job scripts and logs in the results).
Slurm job scripts¶
job-vae-single.sh runs vae-ddp.py as one srun step; job-vae-core-extra.sh runs the core/extra split as two steps. Run either with --help for all options. Their #SBATCH lines target Frontier (-A FUS184, 8 ranks × 7 cores per node); see below for Perlmutter.
sbatch examples/vae/script/job-vae-single.sh --method=1 --num-workers=1
sbatch examples/vae/script/job-vae-single.sh --method=1 --gpudirect --image-scale=2
sbatch examples/vae/script/job-vae-core-extra.sh # split-node: 1 core node, the rest extra
sbatch examples/vae/script/job-vae-core-extra.sh --gpudirect --core-nnodes=2
job-vae-core-extra.sh sets up Slingshot networking for its two steps (see Multiple srun steps). --layout=colocate (both steps on the same nodes) works on Perlmutter only.
Both scripts also run on Perlmutter: they detect the machine (NERSC_HOST) and use 4 ranks per node, with the training ranks seeing all 4 GPUs of their node (NCCL needs that there). Override the Frontier #SBATCH lines when submitting:
sbatch -A <account> -C gpu --gpus-per-node=4 examples/vae/script/job-vae-single.sh --method=1
Multiple srun steps in one job (method=2, cxi)¶
On Slingshot every srun step gets its own VNI (network isolation ID), and two endpoints can only talk on the same VNI. For core and extra running as separate steps, job-vae-core-extra.sh does both of these:
#SBATCH --network=single_node_vni,job_vni:job_vniadds a job-wide VNI to every step (SLINGSHOT_VNIS=<step VNI>,<job VNI>); on Frontier,single_node_vniis also what gives single-node steps a CXI service at all (without itfi_domain()fails with-38).In each task, before Python starts:
export SLINGSHOT_VNIS=${SLINGSHOT_VNIS##*,}. libfabric’s cxi provider uses only the first VNI listed (the step’s own), so without this reads fail withVNI_NOT_FOUND.
Colocating both steps on the same nodes:
|
|
no |
|
|---|---|---|---|
Frontier |
second step fails to launch: |
same failure |
each step has only its own VNI: extra cannot reach core |
Perlmutter |
second step does not start |
works |
extra cannot reach core |
So the script defaults to --layout=split-node; --layout=colocate (with --overlap) is for Perlmutter. Perlmutter’s single-node steps also work without the --network flags.
Troubleshooting: RDMA fails to connect (cxi)¶
If fi_domain() fails with -38 (Function not implemented), the step has no CXI service: add #SBATCH --network=single_node_vni (needed on Frontier for any single-node step, i.e. a -N 1 job or a one-node step inside a larger job). If ranks in different srun steps can’t reach each other (VNI_NOT_FOUND), see Multiple srun steps above.