Environment variables

Read by DDStore itself (the C++ library, pyddstore, cpu_nic_map):

Variable

Default

Effect

DDSTORE_FABRIC

hsn

libfabric provider for method=1/2: hsn (tcp;ofi_rxm) or cxi (native Slingshot; required for GPUDirect RDMA). See libfabric RDMA.

FABRIC_IFACE

auto

Network interface (libfabric domain, e.g. cxi0, hsn0) for method=1/2. Set it to force one; otherwise picked from the rank’s CPU affinity.

DDSTORE_NIC_MAP

unset

Precomputed CPU→NIC map used for that automatic pick instead of a live hwloc query (python3 -m cpu_nic_map --env). The constructor’s nic_map= argument takes priority.

DDSTORE_HANDSHAKE_DIR

./ddstore_hs

method=2 handshake directory when none is given (C++ API; PyDDStore requires handshake_dir, and the examples fill it from this variable). Must be on a shared filesystem.

DDSTORE_HANDSHAKE_TIMEOUT_S

300

Seconds a method=2 extra member’s join() polls for the core group’s record file.

DDSTORE_PROFILE

off

1 turns on get()/get_batch() timing counters, read with get_profile(name). See Performance.

DDSTORE_MAX_READ_BYTES

1073741824 (1 GiB)

method=1/2: largest single fi_read; longer rows are read in pieces (on Perlmutter’s cxi one 5 GB read fails with EMSGSIZE, 2.5 GB works, and the provider doesn’t report the limit). Lowered to the endpoint’s max_msg_size when the provider reports one.

DDSTORE_ALLTOALL_MAX_BYTES

2097152 (2 MiB)

method=0 get_batch(): bytes each rank receives per exchange round. Must be equal on all ranks.

Read by pyddstore.torch (defaults for arguments not given):

Variable

Default

Effect

DDSTORE_METHOD

0

DistDataset’s backend when method= isn’t passed: 0 MPI RMA, 1 libfabric, 2 file-based handshake. (PyDDStore itself takes method= only.)

DDSTORE_BATCH_GET

1

DistDataset.__getitems__ reads a whole batch with one get_batch() per field; 0 reads one sample at a time.

DDSTORE_HANDSHAKE_DIR, DDSTORE_HANDSHAKE_TIMEOUT_S

./ddstore_hs, 300

method=2 directory, and how long DistDatasetReader waits for the core group to publish.

DDSTORE_N_CORE

unset

DistDatasetReader’s number of core ranks when n_core= isn’t passed (the examples default it to 4).

DDSTORE_AFFINITY_WIDTH / DDSTORE_AFFINITY_OFFSET

0 / 0

ThreadDataLoader: pin worker thread i to CPUs [offset + i·width, offset + (i+1)·width) of the process’s affinity; width 0 = no pinning.

Read by the examples (examples/vae/, examples/scripts/, job scripts):

Variable

Default

Effect

DDSTORE_BACKEND

auto

torch.distributed backend for the examples’ DDP setup (nccl, gloo, xccl).

VAE_PROFILE

off

1: vae-ddp.py prints per-epoch fetch vs compute time.

MASTER_PORT

2345

DDP rendezvous port; the core/extra job script gives each step its own.

System settings that matter on Frontier:

Variable

Effect

SLINGSHOT_VNIS

Set by Slurm per step. With --network=job_vni, keep only the last (job-wide) entry before starting Python so separate srun steps can reach each other — see Multiple srun steps.

GPU_MAX_HW_QUEUES

ROCm hardware queues per GPU per process (default 4); raise it if data-loading threads use their own streams — see HIP streams.