DDStore

DDStore logo

Efficient distributed data loading for distributed data-parallel (DDP) training.

Each MPI rank holds a shard of the full dataset in memory. DDStore exposes a global index space so any rank can read any sample via one-sided remote memory access — either MPI RMA (default) or libfabric RDMA — without coordinator synchronization.

  • Batched reads: get_batch() fetches a whole training batch in one call (one-sided RDMA reads in flight together, or an MPI collective for method=0).

  • GPUDirect RDMA: data can live in, and be read straight into, GPU memory (details).

  • PyTorch integration: pyddstore.torch turns any map-style dataset into a distributed one (DistDataset) and provides a thread-based ThreadDataLoader that is safe with MPI and GPU buffers.

  • Thread-safe reads, a profiler for where read time goes, and a split mode (method=2) where a separate job reads data published by another.

DDStore architecture