GPUDirect RDMA¶
add(), get() and get_batch() accept a CUDA/HIP torch.Tensor in place of a NumPy array, so RDMA reads from or writes directly into GPU memory, with no .cpu()/.to(device) copy. Requires method=1 or 2, DDSTORE_FABRIC=cxi, and a CUDA- or ROCm-enabled PyTorch. A GPU tensor with DDSTORE_FABRIC=hsn (the default) or method=0 raises a clear error instead of silently copying through the host.
import torch
data = torch.rand(1024, 64, dtype=torch.float32, device="cuda")
store.add("features", data) # GPU source, no host copy
out = torch.empty((1, 64), dtype=torch.float32, device="cuda")
store.get("features", out, start=2048) # GPU destination, no host copy
add()with a GPU tensor registers your tensor’s own memory; no copy is made. Keep it alive and unmodified untilfree().PyDDStoreholds a reference as a safety net, and adding the same name again with a GPU tensor is rejected. (With NumPy,add()copies and the array can be reused right away.)The device is synchronized before each GPU transfer (
torch.cuda.synchronize(), once perget()/get_batch()call): the NIC writes outside PyTorch’s stream ordering, and without the sync training hit GPU memory faults. Preferget_batch()on the GPU path so this costs one sync per batch, not per sample.init()/update()stay host-only.Whether GPU destinations are faster than host ones depends on the machine: on Frontier they win from ~12.5 KB rows up, on Perlmutter host destinations win at every size (results).
Examples: test/test_gpu_rdma.py, and --gpu-dest/--gpu-source on vae-ddp.py, vae_extra_train.py and vae_core_server.py.