Skip to content

feat(sweep): ✨ split the SRC sweep among the GPUs of a node - #44

Open
robertodr wants to merge 9 commits into
mainfrom
feat/src-multi-gpu
Open

robertodr wants to merge 9 commits into
mainfrom
feat/src-multi-gpu

Conversation

@robertodr

@robertodr robertodr commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

🤖 AI text below 🤖

Splits one SRC sweep among the GPUs of a single node, so that large stacks (the out-of-core reference problem of #42) finish close to G times faster on G GPUs. The mathematics, the random draws and the result do not change. Multi-node is out of scope for now.

src(N, V, M, U, chi_out=2000, device="gpu",
    resources=Resources(devices=4, scratch_dir="/local/scratch"))

Design

One process, one thread per GPU. The earlier MPI/SPMD design (still in docs/superpowers/, to be dropped) is replaced: nothing in it paid off on a single node. The call works as before: it returns the cores, raises errors normally, needs no mpirun, and reads in-memory inputs once.

  • Sketch split. Each GPU owns a contiguous block of the chi_out sketch columns and builds only those columns of every environment, so the environments (the dominant memory) are split too.
    • The left-to-right pass needs no communication.
    • Per site of the right-to-left pass: GPU 0 gathers the sketch, runs the QR and scatters blocks of rows of the output core. Each GPU projects its rows of the next S, which is then gathered on every GPU.
  • Collectives (_group.py) are allgather, gather and scatter, built from peer copies (memcpyPeerAsync) and ordered with the kernels by CUDA events. If one GPU fails, the group is aborted, the other GPUs are woken and the original error is raised, so nothing hangs.
  • Cores are read once by one loader thread, in sweep order. With peer access, each GPU uploads 1/G of the core and gathers the rest over NVLink/P2P. Without it, each GPU uploads the whole core.
  • Planner: each GPU is planned for its block of the sketch columns. gpu_memory is per GPU (detected: the smallest free memory); host memory and scratch disk are shared.
  • Same draws, same result: the Gaussian draws are shared, so the result equals one GPU up to rounding. With one device, the output is bit-identical to main, including with cutoff and spilling.
  • CPU testing: on the CPU, devices=n simulates the GPUs with threads that share the host, through the same code. This lets the suite test the distributed logic without GPUs.

What is not split (sets the efficiency)

  • Every GPU holds a whole site: the cores, S and the output core. The largest D_M is therefore that of one GPU, about 10⁴ at chi_out = 2000 on an H200.
  • GPU 0 runs the QR while the others wait: 0.1–0.2 s per site at the reference size.
  • Every GPU needs a copy of every core: an NVLink allgather with peer access, a full upload over PCIe without.

Contractions scale as χ²D², these costs at most as D², so large problems should scale close to linearly; small ones will not.

Changes

Area Files
Device helpers: selection, peer access, peer copies utils/_backend.py
Device group and collectives _group.py (new)
Per-device plans, Resources.devices _plan.py
Cores read once in sweep order and shared _sites.py
The split sweep _sweep.py
scaling command (--devices 1 2 4 8), --devices on run, --device cpu for smoke tests benches/large/
"Several GPUs" section, testing notes, README docs/content/docs/, README.md

Testing

  • CPU, in the PR suite:
    • collectives, uneven blocks, many rounds in a row, abort handling;
    • sites shared by the devices, and a failed read reaching every device;
    • simulated devices against one device for 2 and 3 devices, chi_out not a multiple of the device count or smaller than it, a custom device order, cutoff, spilling to disk with cleanup, and an error on one device;
    • per-device plans, and the sharing of host memory and disk.
  • GPU, skipped with fewer than two GPUs:
    • two GPUs against one, with split and full uploads, reaching every environment tier;
    • inputs already on a GPU;
    • an error on one GPU, after which the pool limits are restored;
    • an out-of-range device count.

Status

  • Implementation, CPU tests, docs, scaling benchmark
  • GPU tests on a multi-GPU node: the peer copies, the cross-device events and CuPy's per-thread device switching have not run on hardware yet
  • Scaling runs: 1/2/4 GPUs on Leonardo (4×A100), 1/2/4/8 GPUs on an H100, H200 or GH200 node; results in benches/large/README.md
  • Drop docs/superpowers/

Left for later, if measurements call for them: a distributed CholeskyQR2 instead of the QR on GPU 0, splitting the bond of M for D_M above about 10⁴, and NUMA-local host buffers on GH200 nodes.

@robertodr
robertodr added this pull request to stack #43 September 27, 2026 11:36
@Panadestein
Panadestein force-pushed the feat/src-multi-gpu branch 2 times, most recently from 44cb31e to a5ec1e2 Compare October 5, 2026 15:13
Base automatically changed from feat/src-out-of-core to main October 7, 2026 11:50
robertodr and others added 4 commits October 8, 2026 15:42
Phase 2 of the multi-GPU plan: one process per GPU, NCCL collectives,
the sketch index split among the GPUs, with a communicator passed in
Resources. Targets strong scaling at the Phase 1 reference size.

Assisted-by: Pi:claude-opus-5-5
An abort method on the communicator, the truncation rank agreed through
a callback, output sinks, a logged warning for large in-memory inputs,
and MPICH in CI and in the dev shell.

Assisted-by: Pi:claude-opus-5-5
Eleven tasks, each test-first, from the communicator to the scaling
runs. Written under the no-full-code rule of the design phase:
interfaces, pseudocode and test specifications.

Assisted-by: Pi:claude-opus-5-5
Phase 1 is merged, logging is the standard library, the docs moved to
the Fumadocs site, and the CUDA 13 CuPy needs the cu13 NCCL wheel.

Assisted-by: Pi:claude-opus-5.5
Allgather, gather and scatter over peer copies, ordered by events, in a
single process. A failing rank aborts the group so that no rank waits
forever; on the CPU the devices are simulated for testing.

Assisted-by: Pi:claude-opus-5.5
Resources(devices=...) gives each device a contiguous block of the
sketch columns. The left-to-right pass needs no communication; per site
of the right-to-left pass the sketch is gathered on the first device for
the QR, the output core scattered by rows, and the projected environment
gathered on every device. Cores are read once and, with peer access,
uploaded in slices and gathered over the peer links. One device keeps
the previous sweep, bit for bit.

Assisted-by: Pi:claude-opus-5.5
Assisted-by: Pi:claude-opus-5.5
@Panadestein Panadestein changed the title feat(sweep): ✨ SRC sweep over the GPUs of a node feat(sweep): ✨ split the SRC sweep among the GPUs of a node Oct 8, 2026
@Panadestein
Panadestein marked this pull request as ready for review October 8, 2026 14:48
@github-actions

github-actions Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Test Results

    5 files  ±  0      5 suites  ±0   3m 15s ⏱️ +16s
  262 tests + 56    256 ✅ + 51   6 💤 + 5  0 ❌ ±0 
1 251 runs  +265  1 238 ✅ +255  13 💤 +10  0 ❌ ±0 

Results for commit 7e7bdb9. ± Comparison against base commit 698169a.

This pull request removes 2 and adds 58 tests. Note that renamed tests count towards both.
tests.test_sites ‑ test_sites_are_read_when_requested[0]
tests.test_sites ‑ test_sites_are_read_when_requested[1]
tests.test_backend ‑ test_copy_into_on_host
tests.test_backend ‑ test_copy_into_rejects_mismatches
tests.test_backend ‑ test_host_devices_are_simulated
tests.test_gpu_backend ‑ test_too_many_devices_raise
tests.test_gpu_backend ‑ test_two_gpus_match_one[full]
tests.test_gpu_backend ‑ test_two_gpus_match_one[split]
tests.test_gpu_backend ‑ test_two_gpus_raise_the_first_error
tests.test_gpu_backend ‑ test_two_gpus_take_device_inputs
tests.test_group ‑ test_a_failing_rank_releases_the_others
tests.test_group ‑ test_a_group_of_one_runs_inline
…

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants