Performance¶
Use batched reads (the default with
DistDataset, orget_batch()directly). They cut per-sample cost by 10–27× for small rows and make the GPU path insensitive to worker threads; in the VAE every configuration got 1.2–3.9× faster per epoch.Reuse destination buffers and
register_recv()them. Reading into a fresh buffer every time re-registers memory on every read;get_profile(name)["mr_miss"]counts those registrations.method=1(one-sidedfi_read) is the fastest backend;method=0with batching (collective) comes close for small rows.DDSTORE_PROFILE=1+get_profile(name)shows whereget()/get_batch()time goes: lock wait, memory registration, posting and completingfi_read, GPU sync.vae-ddp.pyprints an all-rank summary when it is set. examples/scripts/bench_get.py measures per-row latency and throughput vs row size, destination, batch size and threads.
Measurements, profiles and the experiments behind these choices: docs/results.md.