Repository navigation
Conversation
robertodr
added this pull request to stack #43
September 27, 2026 11:36
robertodr
force-pushed
the
feat/src-multi-gpu
branch
from
September 30, 2026 19:55
fb7dca8 to
e173b01
Compare
Panadestein
force-pushed
the
feat/src-multi-gpu
branch
2 times, most recently
from
October 5, 2026 15:13
44cb31e to
a5ec1e2
Compare
Phase 2 of the multi-GPU plan: one process per GPU, NCCL collectives, the sketch index split among the GPUs, with a communicator passed in Resources. Targets strong scaling at the Phase 1 reference size. Assisted-by: Pi:claude-opus-5-5
An abort method on the communicator, the truncation rank agreed through a callback, output sinks, a logged warning for large in-memory inputs, and MPICH in CI and in the dev shell. Assisted-by: Pi:claude-opus-5-5
Eleven tasks, each test-first, from the communicator to the scaling runs. Written under the no-full-code rule of the design phase: interfaces, pseudocode and test specifications. Assisted-by: Pi:claude-opus-5-5
Phase 1 is merged, logging is the standard library, the docs moved to the Fumadocs site, and the CUDA 13 CuPy needs the cu13 NCCL wheel. Assisted-by: Pi:claude-opus-5.5
Panadestein
force-pushed
the
feat/src-multi-gpu
branch
from
October 8, 2026 14:03
a5ec1e2 to
03d26a6
Compare
Assisted-by: Pi:claude-opus-5.5
Allgather, gather and scatter over peer copies, ordered by events, in a single process. A failing rank aborts the group so that no rank waits forever; on the CPU the devices are simulated for testing. Assisted-by: Pi:claude-opus-5.5
Resources(devices=...) gives each device a contiguous block of the sketch columns. The left-to-right pass needs no communication; per site of the right-to-left pass the sketch is gathered on the first device for the QR, the output core scattered by rows, and the projected environment gathered on every device. Cores are read once and, with peer access, uploaded in slices and gathered over the peer links. One device keeps the previous sweep, bit for bit. Assisted-by: Pi:claude-opus-5.5
Assisted-by: Pi:claude-opus-5.5
Assisted-by: Pi:claude-opus-5.5
Panadestein
marked this pull request as ready for review
October 8, 2026 14:48
Contributor
Test Results 5 files ± 0 5 suites ±0 3m 15s ⏱️ +16s Results for commit 7e7bdb9. ± Comparison against base commit 698169a. This pull request removes 2 and adds 58 tests. Note that renamed tests count towards both. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Splits one SRC sweep among the GPUs of a single node, so that large stacks (the out-of-core reference problem of #42) finish close to
Gtimes faster onGGPUs. The mathematics, the random draws and the result do not change. Multi-node is out of scope for now.Design
One process, one thread per GPU. The earlier MPI/SPMD design (still in
docs/superpowers/, to be dropped) is replaced: nothing in it paid off on a single node. The call works as before: it returns the cores, raises errors normally, needs nompirun, and reads in-memory inputs once.chi_outsketch columns and builds only those columns of every environment, so the environments (the dominant memory) are split too.S, which is then gathered on every GPU._group.py) are allgather, gather and scatter, built from peer copies (memcpyPeerAsync) and ordered with the kernels by CUDA events. If one GPU fails, the group is aborted, the other GPUs are woken and the original error is raised, so nothing hangs.gpu_memoryis per GPU (detected: the smallest free memory); host memory and scratch disk are shared.main, including withcutoffand spilling.devices=nsimulates the GPUs with threads that share the host, through the same code. This lets the suite test the distributed logic without GPUs.What is not split (sets the efficiency)
Sand the output core. The largest D_M is therefore that of one GPU, about 10⁴ atchi_out = 2000on an H200.Contractions scale as χ²D², these costs at most as D², so large problems should scale close to linearly; small ones will not.
Changes
utils/_backend.py_group.py(new)Resources.devices_plan.py_sites.py_sweep.pyscalingcommand (--devices 1 2 4 8),--devicesonrun,--device cpufor smoke testsbenches/large/docs/content/docs/,README.mdTesting
chi_outnot a multiple of the device count or smaller than it, a custom device order,cutoff, spilling to disk with cleanup, and an error on one device;Status
benches/large/README.mddocs/superpowers/Left for later, if measurements call for them: a distributed CholeskyQR2 instead of the QR on GPU 0, splitting the bond of
Mfor D_M above about 10⁴, and NUMA-local host buffers on GH200 nodes.