Skip to content

feat(sweep): ✨ SRC sweep over the GPUs of a node - #44

Draft
robertodr wants to merge 17 commits into
mainfrom
feat/src-multi-gpu
Draft

robertodr wants to merge 17 commits into
mainfrom
feat/src-multi-gpu

Conversation

@robertodr

@robertodr robertodr commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

🤖 AI text below 🤖

Phase 2 of running SRC on stacks too large for one GPU: split the sketch index among the GPUs of a node, so that the Phase 1 reference problem finishes close to G times faster on G GPUs. The mathematics, the random draws and the result stay the same.

Stacked on #42. This draft holds the design and the implementation plan so far. The code will follow on this branch.

Design

Spec: docs/superpowers/specs/2026-09-25-src-multi-gpu-design.md

Question Decision
Execution model one process per GPU (SPMD), launched with srun or mpirun
Collectives NCCL through cupy.cuda.nccl; mpi4py for bootstrap
Public interface an mpi4py communicator in Resources(comm=...); nothing is collective without it
Output rank 0 in memory by default; with Resources(output_dir=...), .npy files and memmaps on every rank
Large inputs file-backed, shared among the ranks of a node through the page cache
Success criterion strong scaling: at least 80% parallel efficiency at 8 GPUs on 8xA100 and 8xH100 at the reference size; at least 90% on 2xA100
  • Communicator: a small interface in _comm.py with three implementations:
    • single process;
    • mpi4py on host arrays, so CI can test the distributed logic without GPUs;
    • NCCL.
  • Driver: each rank owns a cyclic subset of the sketch columns. Per site of the right-to-left pass it does two all-gathers, one of the sketch columns and one of the rows of the projected environment. The QR is repeated on every rank.
  • Errors: errors before the sweep starts are raised on every rank; an error during the sweep aborts the job.
  • Future extensions: the spec records GPUDirect Storage and compression as options for later.

Status

  • Design (spec)
  • Implementation plan (docs/superpowers/plans/2026-09-25-src-multi-gpu.md)
  • Implementation, with unit tests and MPI tests on the CPU
  • Multi-GPU tests and cluster scaling runs

Phase 1 of the multi-GPU plan: stream the input cores, batch the
contractions to a memory budget and keep the environments on the GPU,
in host memory or on local disk. Targets N.V.M.U with D_M = 4000 and
chi_out = 2000 in complex128 on one A100-40GB.

Assisted-by: Pi:claude-opus-5-5
Eleven tasks from the backend helpers to the cluster acceptance run,
each test-first. The code was prototyped and every intermediate state
passes the CPU suite; five refinements of the spec are listed up front.

Assisted-by: Pi:claude-opus-5-5
Host-to-device copies on the compute stream, batch sizes by binary
search over the walked path, the promoted working dtype, lazy sites in
bra stacks, and tests on dense operators rather than cores.

Assisted-by: Pi:claude-opus-5-5
Phase 2 of the multi-GPU plan: one process per GPU, NCCL collectives,
the sketch index split among the GPUs, with a communicator passed in
Resources. Targets strong scaling at the Phase 1 reference size.

Assisted-by: Pi:claude-opus-5-5
An abort method on the communicator, the truncation rank agreed through
a callback, output sinks, a logged warning for large in-memory inputs,
and MPICH in CI and in the dev shell.

Assisted-by: Pi:claude-opus-5-5
Eleven tasks, each test-first, from the communicator to the scaling
runs. Written under the no-full-code rule of the design phase:
interfaces, pseudocode and test specifications.

Assisted-by: Pi:claude-opus-5-5
Base automatically changed from feat/src-out-of-core to main October 7, 2026 11:50

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant