DDStore¶
Efficient distributed data loading for distributed data-parallel (DDP) training.
Each MPI rank holds a shard of the full dataset in memory. DDStore exposes a global index space so any rank can read any sample via one-sided remote memory access — either MPI RMA (default) or libfabric RDMA — without coordinator synchronization.
Batched reads:
get_batch()fetches a whole training batch in one call (one-sided RDMA reads in flight together, or an MPI collective formethod=0).GPUDirect RDMA: data can live in, and be read straight into, GPU memory (details).
PyTorch integration:
pyddstore.torchturns any map-style dataset into a distributed one (DistDataset) and provides a thread-basedThreadDataLoaderthat is safe with MPI and GPU buffers.Thread-safe reads, a profiler for where read time goes, and a split mode (
method=2) where a separate job reads data published by another.
Getting started
User guide
Reference
PyDDStorereferencePyDDStore(comm_or_none=None, method=0, handshake_dir="", n_core=0, nic_map=None)init(name, nrows, disp, itemsize=1)add(name, arr)update(name, arr, offset)get(name, arr, start=0)get_batch(name, arr, indices)register_recv(name, arr)/unregister_recv(name, arr)get_profile(name)join(name)info(name)epoch_begin()/epoch_end()free()
pyddstore.torchreference- Environment variables