Testing¶
Unit tests (pytest)¶
Install test dependencies:
pip install pytest pytest-mpi
Single-rank — no cluster required, covers all dtypes, add/get/init/update, and error cases:
mpirun -n 1 python -m pytest test/test_single.py -v
Multi-rank — verifies remote reads across all rank pairs and sub-communicator grouping:
mpirun -n 4 python -m pytest test/test_multirank.py -v
Batched reads and the PyTorch layer — method 0 everywhere; method 1, method 2 and GPU cases run inside a Slurm step with a CXI device (provider from DDSTORE_FABRIC, default cxi):
mpirun -n 4 python -m pytest test/test_get_batch.py test/test_torch.py -v
GPUDirect RDMA — requires a live cxi fabric and a CUDA/HIP GPU per rank (skipped automatically otherwise); see GPUDirect RDMA:
DDSTORE_FABRIC=cxi mpirun -n 2 python -m pytest test/test_gpu_rdma.py -v
Test file |
Min ranks |
What is tested |
|---|---|---|
|
1 |
All dtypes, |
|
2 (4 recommended) |
Remote reads, shard boundaries, multiple variables, |
|
2 |
GPU-resident |
|
2 (4 recommended) |
|
|
2 (4 recommended) |
|
Integration scripts¶
# Basic functional test (libfabric, method=1)
mpirun -n 4 python examples/scripts/demo.py
# Integration test with PyTorch DDP (libfabric, method=1)
mpirun -n 4 python examples/scripts/test.py
Optional arguments for examples/scripts/demo.py and examples/scripts/test.py:
Flag |
Default |
Description |
|---|---|---|
|
|
Rows per rank |
|
|
Elements per row |
|
|
Number of random reads |
|
|
|
Method 2 (file-based handshake)¶
Two separate launches sharing a handshake directory on a shared filesystem — not a single mpirun, since core and extra are independent jobs:
# Terminal 1 — core (data-holding) side
mpirun -n 4 python test/test_method2_core.py /path/to/shared/ddstore_hs
# Terminal 2 — extra (reader) side, after or while the core side is running
python test/test_method2_extra.py /path/to/shared/ddstore_hs 4