Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
19 commits
Select commit Hold shift + click to select a range
9114854
docs(spec): 📝 design for an out-of-core SRC sweep on one GPU
robertodr Sep 25, 2026
cc951eb
docs(plan): 📝 implementation plan for the out-of-core SRC sweep
robertodr Sep 25, 2026
3787427
docs(spec): 📝 fold the refinements found while planning into the spec
robertodr Sep 25, 2026
bfcf293
feat(utils): ✨ stream, staging and memory helpers for both backends
robertodr Sep 25, 2026
c7508c3
feat(tensor-train): ✨ per-site padding and lazily read sites
robertodr Sep 25, 2026
ddbef8d
refactor(sweep): ♻️ batched site kernels with a peak-memory estimate
robertodr Sep 25, 2026
0a16956
feat(plan): ✨ Resources and memory budget detection
robertodr Sep 25, 2026
2a2848b
feat(plan): ✨ plan batches and environment tiers to the budgets
robertodr Sep 25, 2026
0cd6879
feat(store): ✨ keep environments on the device, in host memory or on …
robertodr Sep 25, 2026
410b7fc
feat(sites): ✨ read the cores of a stack one site at a time
robertodr Sep 25, 2026
55760b6
feat(sweep): ✨ out-of-core SRC sweep planned to memory budgets
robertodr Sep 25, 2026
a5d0693
test(gpu): ✅ every environment tier and the pool cap on the GPU
robertodr Sep 25, 2026
ac13e40
perf(bench): 📈 out-of-core acceptance benchmark for N.V.M.U
robertodr Sep 25, 2026
2bfa815
docs: 📝 large problems, Resources and the out-of-core test knobs
robertodr Sep 25, 2026
8b5a48b
chore: 🔀 merge main into feat/src-out-of-core
Panadestein Oct 7, 2026
8a7ce2c
fix(plan): 🐛 cap the GPU pool at the budget plus the margin
Panadestein Oct 7, 2026
da39b0e
chore: 🔥 remove the out-of-core implementation plan
Panadestein Oct 7, 2026
b9216fd
chore: 🔥 remove the out-of-core design spec
Panadestein Oct 7, 2026
9d0563f
chore(deps): 🔧 move the NVIDIA extra to CUDA 13
Panadestein Oct 7, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 21 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,26 @@ The result then pairs with a ket by plain contraction, with no further conjugati

At runtime, pass `device="gpu"` to use GPU acceleration. The library handles backend dispatch automatically.

### Large Problems

Stacks that do not fit on the GPU, or in host memory, still run: the sweep reads the input cores one site at a time (from NumPy arrays, `np.memmap`, zarr or HDF5 datasets), batches every contraction to a memory budget, and keeps the sketched environments on the GPU, in host memory or on local disk. The budgets are detected, or set explicitly with `Resources`:

```python
from src_method import Resources, src

out = src(
N,
V,
M,
U,
chi_out=2000,
device="gpu",
resources=Resources(gpu_memory="36GB", scratch_dir="/local/scratch"),
)
```

See [Large problems](https://algorithmiq.github.io/src-method/features/large-problems) for the budgets, the scratch directory and how to read the logged plan.

### Logging

`src_method` logs through the standard library `logging` module, under the
Expand All @@ -108,7 +128,7 @@ logging.getLogger("src_method").setLevel(logging.DEBUG)
# CPU only (default)
uv pip install src_method

# With NVIDIA GPU support (CUDA 12.x)
# With NVIDIA GPU support (CUDA 13.x, driver >= 580)
uv pip install "src_method[gpu-nvidia]"

# With AMD GPU support (ROCm)
Expand Down
3 changes: 3 additions & 0 deletions benches/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,9 @@ different SRC variants. They are grouped by the kind of workload exercised:
- [`stack/`](stack/) — One-shot SRC over stacks of trains against sequential
pairwise application, by stack depth: accuracy against dense references and
wall time. See [`stack/README.md`](stack/README.md).
- [`large/`](large/) — Out-of-core SRC of `N . V . M . U` with a large `M` read
from disk, on one GPU: plan, wall time, memory peaks and stall time. See
[`large/README.md`](large/README.md).

The scripts need the `bench` dependency group (included in `dev`):

Expand Down
33 changes: 33 additions & 0 deletions benches/large/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
# Out-of-core benchmark

`bench_large.py` compresses the stack `N . V . M . U` of Pauli-transfer-matrix MPOs
(physical legs of 4) on one GPU, where `M` has a large bond and is read from disk.

```bash
# Write the stack to node-local NVMe (205 GB for M at the defaults, complex128).
uv run python benches/large/bench_large.py generate /local/stack --bond-m 4000

# Compress it with detected budgets, spilling to the same disk.
uv run python benches/large/bench_large.py run /local/stack --chi-out 2000 \
--scratch-dir /local/scratch

# Check that the plan does not change the result: two GPU budgets, same seed.
uv run python benches/large/bench_large.py generate /local/medium --bond-m 1000
uv run python benches/large/bench_large.py compare /local/medium --chi-out 500 \
--small 40GB --large 80GB --scratch-dir /local/scratch
```

`run` logs the wall time, the size of CuPy's pool at the end (its high-water mark,
since the pool keeps its blocks), the host peak (`ru_maxrss`), the planned peaks,
the bytes spilled and the tiers. With `--debug` before the command, `src_method` also
logs its plan, the time of each pass and the stall time (`SRC stalls`).

Acceptance of the out-of-core sweep on one A100-40GB, with at least 300 GB of host
memory and node-local NVMe:

1. `run` completes at 50 sites, `D_M = 4000`, `chi_out = 2000`, complex128.
2. The pool size stays within the GPU budget plus the `max(10%, 1 GiB)` margin
the pool cap allows for fragmentation, and the host peak within the host
budget.
3. The stall time is below 10% of the wall time.
4. `compare` at `D_M = 1000`, `chi_out = 500` reports a distance below `1e-10`.
215 changes: 215 additions & 0 deletions benches/large/bench_large.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,215 @@
"""Out-of-core SRC of ``N . V . M . U`` with a large ``M``, read from disk.

Three commands:

- ``generate``: write the four MPOs site by site as ``.npy`` files, so that ``M``
is never whole in memory.
- ``run``: compress the stack with automatic (or given) budgets and report the
plan, the wall time, the device pool size, the host peak and the stall time.
- ``compare``: run twice with two GPU budgets and report the relative distance
between the two outputs, which checks that the plan does not change the result.

The plan, the per-pass times and the stall time are also logged by `src_method`
itself at ``DEBUG``; pass ``--debug`` to see them.
"""

from __future__ import annotations

import logging
import resource
from pathlib import Path # noqa: TC003 (cyclopts reads the annotations at runtime)
from time import perf_counter
from typing import Annotated

import cyclopts
import numpy as np

import src_method._sweep as sweep_module
from src_method import Resources, src
from src_method._plan import make_plan

logger = logging.getLogger(__name__)
app = cyclopts.App(help="Benchmark out-of-core SRC of N.V.M.U with a large M.")

LAYERS = ("N", "V", "M", "U")
PHYS = 4 # Pauli transfer matrix legs


def _site_shape(j: int, n_sites: int, bond: int) -> tuple[int, ...]:
left = () if j == 0 else (bond,)
right = () if j == n_sites - 1 else (bond,)
return (*left, *right, PHYS, PHYS)


@app.command
def generate(
directory: Path,
*,
n_sites: int = 50,
bond: int = 4,
bond_m: int = 4000,
dtype: str = "complex128",
seed: int = 0,
) -> None:
"""Write random MPOs ``N``, ``V``, ``M``, ``U`` as one ``.npy`` file per site.

Args:
directory: Where to write, ideally node-local NVMe.
n_sites: Number of sites.
bond: Bond dimension of ``N``, ``V`` and ``U``.
bond_m: Bond dimension of ``M``.
dtype: ``float64`` or ``complex128``.
seed: Seed of the random draws.
"""
rng = np.random.default_rng(seed)
kind = np.dtype(dtype)
for name in LAYERS:
chi = bond_m if name == "M" else bond
(directory / name).mkdir(parents=True, exist_ok=True)
# Scaled so that products of the layers stay of order one.
scale = 1 / np.sqrt(chi * PHYS)
for j in range(n_sites):
shape = _site_shape(j, n_sites, chi)
site = np.lib.format.open_memmap(
directory / name / f"{j:04d}.npy", mode="w+", dtype=kind, shape=shape
)
for row in range(shape[0]): # one left-bond slice at a time
draw = rng.normal(size=shape[1:])
if kind.kind == "c":
draw = draw + 1j * rng.normal(size=shape[1:])
site[row] = draw * scale
site.flush()
del site
logger.info("Layer %s written: bond %d", name, chi)


def _load(directory: Path) -> list[list[np.ndarray]]:
return [
[
np.load(path, mmap_mode="r")
for path in sorted((directory / name).glob("*.npy"))
]
for name in LAYERS
]


def _run(
directory: Path, chi_out: int, resources: Resources, seed: int
) -> list[np.ndarray]:
layers = _load(directory)
plans = []

def spy(*args: object, **kwargs: object) -> object:
plans.append(make_plan(*args, **kwargs))
return plans[-1]

sweep_module.make_plan = spy # record the plan of the run
try:
start = perf_counter()
out = src(
*layers,
chi_out=chi_out,
dtype=layers[0][0].dtype,
seed=seed,
device="gpu",
resources=resources,
)
seconds = perf_counter() - start
finally:
sweep_module.make_plan = make_plan
import cupy # noqa: PLC0415 (GPU-only benchmark)

(plan,) = plans
tiers = [site.tier for site in plan.sites]
logger.info(
"Run complete: %.1f s, pool %d B, host peak %d B, planned device peak %d B, "
"planned host peak %d B, disk %d B, tiers %s, sketch batches %s",
seconds,
cupy.get_default_memory_pool().total_bytes(),
resource.getrusage(resource.RUSAGE_SELF).ru_maxrss * 1024,
plan.device_peak,
plan.host_peak,
plan.disk_bytes,
{tier: tiers.count(tier) for tier in ("device", "host", "disk")},
sorted({site.sketch_batch for site in plan.sites[1:]}),
)
return out


@app.command
def run(
directory: Path,
*,
chi_out: int = 2000,
gpu_memory: str | None = None,
host_memory: str | None = None,
scratch_dir: Path | None = None,
seed: int = 0,
) -> None:
"""Compress the stack written by ``generate`` on the GPU.

Args:
directory: The directory given to ``generate``.
chi_out: The output bond dimension.
gpu_memory: GPU budget, e.g. ``36GB``; detected when omitted.
host_memory: Host budget; detected when omitted.
scratch_dir: Where to spill environments; ``$TMPDIR`` when omitted.
seed: Seed of the sketch.
"""
_run(directory, chi_out, Resources(gpu_memory, host_memory, scratch_dir), seed)


def _inner(a: list[np.ndarray], b: list[np.ndarray]) -> complex:
"""Frobenius inner product ``<a, b>`` of two MPOs, site by site."""
env = np.einsum("rud,sud->rs", a[0].conj(), b[0])
for x, y in zip(a[1:-1], b[1:-1]):
env = np.einsum("rs,rtud,svud->tv", env, x.conj(), y, optimize=True)
return complex(np.einsum("rs,rud,sud->", env, a[-1].conj(), b[-1]))


@app.command
def compare(
directory: Path,
*,
chi_out: int = 500,
small: str = "40GB",
large: str = "80GB",
scratch_dir: Path | None = None,
seed: int = 0,
) -> None:
"""Run with two GPU budgets and report the relative distance of the outputs.

Args:
directory: The directory given to ``generate``.
chi_out: The output bond dimension.
small: The smaller GPU budget.
large: The larger GPU budget.
scratch_dir: Where to spill environments; ``$TMPDIR`` when omitted.
seed: Seed of the sketch, the same for both runs.
"""
first = _run(directory, chi_out, Resources(small, None, scratch_dir), seed)
second = _run(directory, chi_out, Resources(large, None, scratch_dir), seed)
aa, bb, ab = _inner(first, first), _inner(second, second), _inner(first, second)
distance = np.sqrt(max((aa + bb - 2 * ab).real, 0.0) / aa.real)
logger.info("Relative distance between the runs: %.3e", distance)


@app.meta.default
def main(
*tokens: Annotated[str, cyclopts.Parameter(show=False, allow_leading_hyphen=True)],
debug: bool = False,
) -> None:
"""Configure logging, then dispatch to a command.

Args:
tokens: The command and its arguments.
debug: Also show the `src_method` debug log: plan, pass times and stalls.
"""
logging.basicConfig(level=logging.INFO, format="%(message)s")
if debug:
logging.getLogger("src_method").setLevel(logging.DEBUG)
app(tokens)


if __name__ == "__main__":
app.meta()
18 changes: 18 additions & 0 deletions docs/content/docs/contributing/testing.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,11 @@ Always run tests through `pytest`, never as `python tests/test_<name>.py`.
| `test_package.py` | `apply` and `compress` against `quimb` references, validation, precision |
| `test_stack.py` | `src` over stacks and the sweep behind it |
| `test_tensor_train.py` | the shared train helpers in `_tensor_train` |
| `test_backend.py` | the stream, staging and memory helpers in `utils._backend` |
| `test_kernels.py` | the batched site contractions and their peak-memory estimate |
| `test_plan.py` | the planner: budgets, batch sizes and environment tiers |
| `test_store.py` | environment storage on the device, in host memory and on disk |
| `test_sites.py` | reading the cores one site at a time, with prefetching |
| `test_logging.py` | the package leaves the host's logging configuration alone |
| `test_gpu_backend.py` | the CuPy path; skipped without CuPy or a GPU |

Expand Down Expand Up @@ -70,6 +75,19 @@ is the behaviour under test, assert it with `pytest.warns`. The exact fallback f
networks with fewer than three sites logs a warning rather than raising one, so
tests of that path assert on the log instead, as below.

## Exercising the out-of-core paths

The batched contractions and the host and disk tiers of the sweep only engage when
the budgets are tight, so tests force them with tiny budgets: on the CPU,
`Resources(host_memory="1MB", scratch_dir=tmp_path)` gives small batches and spills
the environments of a depth-4 stack with bonds of 4 to `tmp_path` (see
`tests/test_stack.py`). Compare the dense operator of the result with that of a
default run, not the cores: batching changes the rounding, and with it the cores of
an ill-conditioned sketch, but not the operator they represent. The planner is a
pure function, so `tests/test_plan.py` checks batch sizes and tiers from shapes
alone, and `tests/test_gpu_backend.py` derives GPU budgets from a plan made with
`make_plan` to reach every tier.

## Logging in tests

`log_cli` is on, so log records show up live while the tests run. To assert on
Expand Down
2 changes: 1 addition & 1 deletion docs/content/docs/features/gpu.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,7 @@ through [CuPy](https://cupy.dev/). Both NVIDIA (CUDA) and AMD (ROCm) GPUs are
supported through optional extras:

```bash
uv pip install "src_method[gpu-nvidia]" # CUDA 12.x
uv pip install "src_method[gpu-nvidia]" # CUDA 13.x, driver >= 580
uv pip install "src_method[gpu-rocm]" # ROCm
```

Expand Down
3 changes: 2 additions & 1 deletion docs/content/docs/features/index.mdx
Original file line number Diff line number Diff line change
@@ -1,12 +1,13 @@
---
title: Features
description: Stacks, adaptive truncation, precision, GPUs and logging.
description: Stacks, adaptive truncation, precision, GPUs, large problems and logging.
---

<Cards>
<Card title="Stacks" href="/features/stacks" description="Contract and compress several trains in a single sweep." />
<Card title="Adaptive truncation" href="/features/cutoff" description="Trim bonds below chi_out with a relative singular-value cutoff." />
<Card title="Precision and reproducibility" href="/features/precision" description="Data types of sketches and results, and seeding." />
<Card title="GPU execution" href="/features/gpu" description="Run the sweep on NVIDIA or AMD GPUs through CuPy." />
<Card title="Large problems" href="/features/large-problems" description="Run stacks that fit neither on the GPU nor in host memory." />
<Card title="Logging" href="/features/logging" description="Turn on the library's progress and timing messages." />
</Cards>
Loading
Loading