# Devices

A run executes on one device kind, chosen at start with `--device` or
`SEXPGPU_DEVICE`. The file never names a device; the same IR runs on every backend.

| device | what it is | for |
|---|---|---|
| `cpu` | the IR interpreter, the default | compiling, reading data, differentiating and real steps at toy sizes, on any machine |
| `metal` | generated Metal compute kernels on Apple silicon | local training and evaluation, at toy and small sizes today |
| `cuda` | the CUDA backend: fused generated kernels, cuBLAS, and native cuDNN product and attention graphs | training |

## The CPU interpreter

The interpreter keeps every node's value alive for the whole graph and runs
one thread, so it cannot hold an LM-scale shape: a full vocabulary at
sequence 1024 costs an hour and several gigabytes for two steps.

A run at an LM-scale shape therefore declares a `smoke` variant that
shrinks it:

```lisp
(defvariant smoke (layers 2) (width 128) (seq 64))
```

and `sexpgpu run <file> --variant smoke --steps 3 --device cpu` is the
local check before a GPU is paid for. A run without such a variant is
checked on the GPU instead, with `--steps 3`.

## The Metal device

`--device metal` or `SEXPGPU_DEVICE=metal` selects the Apple GPU.
It requires an Apple silicon Mac, macOS 12 or newer, and a binary built
with `--features metal`. Metal compiles kernels at runtime through the
system framework. Training needs neither CUDA nor the Xcode command-line tools.

From the `sexpgpu` directory, with Rust 1.89 or newer:

```bash
cargo build --release -p sexpgpu --features metal
target/release/sexpgpu run crates/cli/tests/fixtures/tiny-sgd.sx --device metal
```

The backend implements every executable IR operation, including gradients,
convolution, scans, table scans and top-k. Training, generation,
evaluation, metrics, and checkpoints use the same loop as the other
devices. Parameters and
optimizer state remain in Metal buffers between calls. Kernels and graph
plans are cached for the life of the process.

The backend uses the shared planner's fusion, tiled SIMD-group matrix
products, parallel reductions, and segmented long scans. One-pass F32 scans
can fuse a short additive expression used only by that scan, avoiding its
intermediate buffer. The recurrence order, explicit casts, and launch
geometry stay unchanged. The Metal internals
describe which expressions qualify, and the
scan-input report records the evidence.

The backend reuses intermediate buffers after their last read. Each device
pool initially keeps at most 256 MiB of idle buffers. An allocation request
above 256 MiB promotes that pool after the request passes the device's
maximum-buffer check. The promoted idle limit is the smaller of 8 GiB and
half the device's recommended
maximum working set, fixed for the pool's remaining lifetime. Live allocations
are outside this idle limit. After promotion, newly returned idle buffers are
made purgeable so macOS can discard their storage under memory pressure.
Buffers cached before promotion can stay resident until reuse, and some
small shared buffers cannot become purgeable. Together their cached lengths
stay bounded by the original 256 MiB allowance. This is not an RSS limit. Reuse
restores nonpurgeable storage before writing new values. Buffers referenced
by pending GPU commands stay nonpurgeable until those commands complete.
If a new allocation fails, the backend releases cached buffers and retries
once. Scatters fold repeated indices in input order.
Each graph waits for its command buffer before returning.

There are no dedicated attention kernels, memory admission model,
automatic stacking, or multi-GPU collectives. Without a memory plan, a
buffer larger than the device's advertised maximum, or an allocation Metal
cannot satisfy, stops the run with `E-MEM-002`. The maximum is checked
before requesting the buffer:

```text
sexpgpu run: E-MEM-002: the Metal device could not allocate 34358689800 bytes for a tensor of [65535, 65535] i64; the run does not fit this GPU
```

On an M1 Pro with 32 GiB, the full `tiny-adam.sx` model with vocabulary 50304
takes about 3.11 seconds per F32 step in two short six-step runs, versus
5.11 seconds with the original 256 MiB pool. The
pool measurements include
complete-step timings and their limits. Large training runs and performance
parity with PyTorch remain unqualified.

BF16 values use F32 storage and round at explicit IR casts, matching the
CPU interpreter. BF16 therefore does not halve memory use on Metal, and its
cast kernels make it no faster.
Integers use I64 storage, including values with IR dtype `i32`.
`SEXPGPU_DEVICES` accepts only `0` or an unset value for Metal.
CUDA-specific kernel switches do not change Metal execution.

## The CUDA device

`SEXPGPU_DEVICE=cuda` needs the Linux CUDA binary, `sexpgpu-linux-cuda`.

| requirement | why |
|---|---|
| a CUDA 13 driver, 580 series or newer | the binary loads the 13.x driver, NVRTC and cuBLAS at start |
| compute capability 8.0 or newer: A100, L4, H100 | bf16 tensor-core GEMMs need Ampere; an older card is refused when the device opens. The native cuDNN graphs are used on `compute_80` only |
| cuDNN 9.13.0 for CUDA 13, optional | the native product and attention graphs; without it those regions run the ordinary lowering, and the [lowering report](https://sx.041.io/docs/explain.md#the-lowering-report) says so |

Nothing else: no Rust, no Python, no checkout. Kernels are compiled at run
time by NVRTC from source inside the binary. [`sexpgpu doctor`](https://sx.041.io/docs/doctor.md)
checks the floor.

What the device did with each graph is the
[lowering report](https://sx.041.io/docs/explain.md#the-lowering-report). `runtime/peak_bytes`
reports the allocator's high-water mark every step; see [metrics](https://sx.041.io/docs/metrics.md).

## Memory

Before its first step a CUDA run works out, from the plan its device will
execute, the most it will hold at once on each GPU, and chooses how many
microbatches one call of the training graph runs (its stacking). Both are
lines of the lowering report; a stacking measured at the first step is
printed on its own once timed. The run ends with the peak it reached
against the plan:

```text
lowering: cuda compute_80 NVIDIA A100-SXM4-80GB, patterns on
  ...
  memory                    2.7 GiB of 78.8 GiB free
  stacking                  measured at the first step
...
stacking 1: measured at the first step, the model within 5% could not separate them (4 130.2 ms, model 130.5 ms; 2 135.2 ms, model 131.5 ms; 1 128.9 ms, model 133.6 ms)
...
memory: peak 2.8 GiB of 2.7 GiB planned (+1.8%)
```

- **Stacking** is chosen among the degrees that fit: the only one, the cost
  model's prediction, a measurement at the first step when the predictions
  are within 5 percent, or a measurement an earlier run of the same graph
  made on the same device, remembered in `~/.cache/sexpgpu/choices.json`
  (delete it to measure again). A measured choice can differ on another
  device, where the F32 products then sum in another order; see
  [determinism](https://sx.041.io/docs/determinism.md).
- **The plan** keeps 1 GiB free for the libraries. It stacks F32
  microbatches only when the stacked graph fits, drops cached parameter
  results when only that fits, and otherwise refuses with `E-MEM-001`
  before any initializer runs. Optimizer diagnostics select an alternate
  update graph. The plan counts one proposal per update, including its
  selected diagnostic outputs, until every update can commit.
- **An allocation that fails anyway** releases the device's caches and
  retries, and says so in a `memory: allocation failed` line and annotation.
  A training microbatch is retried once. When the retry fails too, or when
  the third microbatch in a row cannot allocate on its first try, the run
  stops with `E-MEM-003`. The error names the step, the microbatch, the
  planned memory and what is free on the device. A rank that stops this
  way stops the whole run.
- **A [generator](https://sx.041.io/docs/generate.md)** is a phase of the plan: each of its graphs
  beside the batch and the state its trips carry, stacked with the
  training graph's degree. The `memory` line adds its largest,
  `generation 0.4 GiB`, and the cost model prices a step's generation with
  its training calls.
- **Both lines are annotations** on the metrics stream, the memory one with
  its parts: parameters, optimizer states, gradients, constants, collective
  buffers, the largest graph's live set, a generator's largest phase, and
  where it peaks.
- **`SEXPGPU_MEMORY`** plans against a smaller card than the one present.

Cross-node CUDA runs reserve the aligned gradient transport buffer. BF16
transport also reserves a narrowed copy and F32 error feedback for F32
gradients. These buffers persist across steps. One-node runs reserve none.
The executor's `runtime/peak_bytes` and `held_bytes` exclude these
coordinator-owned allocations; admission includes them under `collective
buffers`, and the final peak line compares against the plan less them.

CUDA also reserves 256 MiB for persistent execution of small graphs, or
4 GiB when a larger graph qualifies. The final
memory line reports that cache's peak and reservation separately from the
graph live-set comparison. `held_bytes` includes its current allocations.
`runtime/peak_bytes` adds the private cache's peak to the ordinary allocator's
high-water mark, so it is a conservative sum of the two peaks.

## Several GPUs

`SEXPGPU_DEVICES=0,1` runs data parallel over those CUDA ordinals, in rank
order; unset uses every visible device, and one ordinal is the single-GPU
path. `CUDA_VISIBLE_DEVICES` limits what is visible.

- The global `defrun :microbatches` is split over the GPUs, so the device
  count must divide it. Each rank reads its own share of the loader.
- The result is the same experiment: one optimizer step per step, gradients
  combined across ranks.
- Timing series and sampled diagnostics are rank zero's; a diagnostic's
  reducer folds across ranks.
- A resume needs the device count the checkpoint was written with.
- `sexpgpu` sets `NCCL_RUNTIME_CONNECT=0` unless it is set, so NCCL
  allocates its transport buffers when the communicators are created,
  before the memory plan reads free memory, instead of at the first
  gradient sum, when the pool may hold the rest of the card.

### Several nodes

One run can span machines: one `run` process on each node, each with
`SEXPGPU_DEVICE=cuda` and the same number of GPUs. A node's GPUs are the
global ranks after the previous nodes'.

| variable | meaning |
|---|---|
| `SEXPGPU_NODES` | the node count; unset or `1` is one node |
| `SEXPGPU_NODE_RANK` | this process's node, `0` to nodes - 1. Node 0 writes the metrics, the status file and the checkpoints, so a checkpoint location every node reads is an `s3://` one |
| `SEXPGPU_RENDEZVOUS` | `host:port` of node 0, where it listens once for the other nodes. Under SkyPilot the host is the first line of `SKYPILOT_NODE_IPS` |
| `SEXPGPU_NODE_GRADIENTS` | `bf16` (default) rounds each device's f32 gradients to bf16 for the sum between nodes, carrying each rounding's error into the next step, half the bytes on the wire; `f32` sends them exactly, so two nodes of one GPU train bitwise as one node of two and a resume is exact, and is what `--deterministic` sends. Every node sets the same |

Under SkyPilot with `num_nodes`, each node's `run` sets the three from
SkyPilot's own variables:

```bash
export SEXPGPU_NODES="$SKYPILOT_NUM_NODES" SEXPGPU_NODE_RANK="$SKYPILOT_NODE_RANK"
export SEXPGPU_RENDEZVOUS="$(echo "$SKYPILOT_NODE_IPS" | head -n1):29500"
```

A resume needs the node count the checkpoint was written with.

### Errors

They stop `run` with exit `1`, before or during training; `doctor` reports
`E-DP-001` and `E-DP-005`.

| code | when | fix |
|---|---|---|
| `E-DP-001` | `SEXPGPU_DEVICES` does not parse, repeats an ordinal, names one that is not visible, or no GPU is visible | list visible ordinals once each, `SEXPGPU_DEVICES=0,1` |
| `E-DP-002` | a resume's rank count, node count or loader rank differs from the checkpoint's; the message names both | resume with the checkpoint's nodes and devices |
| `E-DP-003` | the global microbatches do not divide over the GPUs | pick a device count that divides `:microbatches` |
| `E-DP-004` | the ranks' or the nodes' parameters, optimizer states, or CUDA product choices disagree at a checkpoint | a backend bug, never the experiment's fault: that checkpoint was not written, so `--resume latest` continues from the one before; report the full diagnostic and `SEXPGPU_DEVICES` |
| `E-DP-005` | `SEXPGPU_NODES` is not a count, `SEXPGPU_NODE_RANK` is not in `0..nodes`, `SEXPGPU_RENDEZVOUS` is not `host:port`, `SEXPGPU_NODE_GRADIENTS` is not `f32` or `bf16`, or the device is not `cuda` | set all three on every node, with `SEXPGPU_DEVICE=cuda` |
| `E-DP-006` | a node differs from node 0 at the rendezvous: node count, a node index twice, GPU count, experiment, gradient dtype, or the checkpoint it resumes; the message names both | the same files, flags, device count and checkpoint location on every node |
| `E-DP-007` | a node did not arrive at the rendezvous within 300 s, or left the run | start every node; after a loss, restart every node with `--resume latest` |

For a disagreement between ranks on one node, `E-DP-004` names the checkpoint
step, the expected and actual global ranks, and the first parameter or optimizer
state that differs. It reports the tensor key and either the schema difference
or the first differing element's flat index and values. Floating-point values
include hexadecimal bits, so signed zeros and NaN payloads remain distinguishable.
A disagreement across nodes names the node whose checkpoint digest differs.

Related: [run](https://sx.041.io/docs/run.md), [doctor](https://sx.041.io/docs/doctor.md), [defrun](https://sx.041.io/docs/defrun.md),
[environment](https://sx.041.io/docs/environment.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
