Repository navigation
Conversation
robertodr
added this pull request to stack #43
September 27, 2026 11:36
robertodr
force-pushed
the
feat/src-multi-gpu
branch
from
September 30, 2026 19:55
fb7dca8 to
e173b01
Compare
Panadestein
force-pushed
the
feat/src-multi-gpu
branch
from
October 5, 2026 14:48
e173b01 to
44cb31e
Compare
Phase 1 of the multi-GPU plan: stream the input cores, batch the contractions to a memory budget and keep the environments on the GPU, in host memory or on local disk. Targets N.V.M.U with D_M = 4000 and chi_out = 2000 in complex128 on one A100-40GB. Assisted-by: Pi:claude-opus-5-5
Eleven tasks from the backend helpers to the cluster acceptance run, each test-first. The code was prototyped and every intermediate state passes the CPU suite; five refinements of the spec are listed up front. Assisted-by: Pi:claude-opus-5-5
Host-to-device copies on the compute stream, batch sizes by binary search over the walked path, the promoted working dtype, lazy sites in bra stacks, and tests on dense operators rather than cores. Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
…disk Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Assisted-by: Pi:claude-opus-5-5
Phase 2 of the multi-GPU plan: one process per GPU, NCCL collectives, the sketch index split among the GPUs, with a communicator passed in Resources. Targets strong scaling at the Phase 1 reference size. Assisted-by: Pi:claude-opus-5-5
An abort method on the communicator, the truncation rank agreed through a callback, output sinks, a logged warning for large in-memory inputs, and MPICH in CI and in the dev shell. Assisted-by: Pi:claude-opus-5-5
Eleven tasks, each test-first, from the communicator to the scaling runs. Written under the no-full-code rule of the design phase: interfaces, pseudocode and test specifications. Assisted-by: Pi:claude-opus-5-5
Panadestein
force-pushed
the
feat/src-multi-gpu
branch
from
October 5, 2026 15:13
44cb31e to
a5ec1e2
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Phase 2 of running SRC on stacks too large for one GPU: split the sketch index among the GPUs of a node, so that the Phase 1 reference problem finishes close to
Gtimes faster onGGPUs. The mathematics, the random draws and the result stay the same.Stacked on #42. This draft holds the design and the implementation plan so far. The code will follow on this branch.
Design
Spec:
docs/superpowers/specs/2026-09-25-src-multi-gpu-design.mdsrunormpiruncupy.cuda.nccl;mpi4pyfor bootstrapmpi4pycommunicator inResources(comm=...); nothing is collective without itResources(output_dir=...),.npyfiles and memmaps on every rank_comm.pywith three implementations:mpi4pyon host arrays, so CI can test the distributed logic without GPUs;Status
docs/superpowers/plans/2026-09-25-src-multi-gpu.md)