# Determinism

What a run may rely on about its bits. Every kernel SexpGPU runs sums in a
fixed order: there are no floating-point atomics, the cuDNN engines it
accepts are the deterministic ones, and cuBLAS keeps its default of no
atomics. What can still move the bits is a decision made from the machine
rather than from the file. This page names each of them, where it is
remembered, and the level of reproduction a run gets.

## Which flag, when

| you want | run with |
|---|---|
| the fastest run, on whatever card you get | nothing: the run measures its choices, remembers them in `~/.cache/sexpgpu`, and repeats its own bits on that machine (level 1) |
| a spot job that may resume on another card | `--choices fresh`, or `SEXPGPU_CHOICES=fresh` in the job's environment: a resume chooses again where its checkpoint's choices do not fit or are not offered, and continues at level 3 rather than stopping |
| a resume that keeps the exact bits of the run it continues | nothing, on the same kind of card: a resume pins itself to its checkpoint's choices (level 2), and stops with `E-CHOICE-001` rather than sum in another order if the card cannot take them |
| the same bits as a colleague's run, on your card | `--choices <their checkpoint or experiment.json>` |
| bits that hold with no record at all, for a test or a paper's comparison | `--deterministic`: a few percent slower, nothing measured |

## The three levels

1. **One device, same bits.** The same file, selection and flags, run twice
   on one machine with the same binary, GPUs, driver, libraries and
   `~/.cache/sexpgpu`, give every loss, metric, parameter and optimizer state
   the same bits. Every run has this level. The cache is what makes it true:
   a decision measured once is remembered and taken again, instead of
   measured again and possibly taken otherwise.
2. **Any device, same bits, given the choices.** A run
   [pinned](#pinning-a-run-to-recorded-choices) to another run's recorded
   choices reproduces its bits on another machine, provided that machine
   can execute every choice and decides the rest of the lowering alike
   ([below](#what-the-device-decides)). A run under `--deterministic` has
   this level against every other `--deterministic` run of the file,
   without any record: its choices are fixed rules.
3. **Same distribution.** Without the same choices another machine may
   decide differently, and its sums run in another order. The loss curve
   stays inside the ladder's envelope for its tier, not bit for bit.

The CPU interpreter decides nothing from the machine; a run on it is at
level 1 on any machine with the same platform math library.

Metal makes no measured choices. It uses stacking 1, zero recomputation,
and the IR composition for convolutions. Its checkpoint identity includes
the Apple GPU name, MPS F32 schedule with K at most 1024, batched leading
sums with at least 128 matrices and 65536 elements per product, independent
matrix batching, and macOS build. Independent batching inserts the literal
`independent-batches` between `M*N>=65536` and `macOS` in `choices.device`.
Lowering annotations repeat that identity in `metadata.device` and
`metadata.choices.device`.
The OS supplies the Metal compiler, driver, and MPS implementation.
Repeated graphs and checkpoint resume reproduce bits on the same
configuration. CPU and CUDA comparisons use numerical tolerances; Metal's
BF16 storage follows the CPU interpreter's cast boundaries. Eligible dense
F32 matrix products use MPS with reduced precision disabled; their
floating-point association can differ from the tiled shaders and the
interpreter. Batched and per-matrix MPS calls can also produce different
product bits. Equality measured on one device does not guarantee equality
on another GPU or macOS build, where MPS can use a different schedule.
SIMD-group products, parallel reductions, and segmented scans use fixed
schedules that can round differently from the interpreter's sequential folds.
Batch sums of tiled products retain each product's rounding and the sum's
order, as when the products are materialized separately.

The [lowering report](https://sx.041.io/docs/explain.md#the-lowering-report) of a CUDA run ends
with a `determinism` line that names the run's level and why, and a
`choices` line with what it chose:

```text
  determinism         1: this machine repeats these bits with this binary and ~/.cache/sexpgpu; another is 3, or 2 pinned to these choices with --choices
  choices             stacking 1, recompute 10.99 flops a byte, 17 convolution engines; on compute_80 with 108 multiprocessors, cuBLAS 130101, cuDNN 91300, NVRTC 13.0, patterns on
```

## What a run chooses

Decisions made by measuring, or from what the machine holds at the time.
They are the ones that can differ between two machines with the same card,
and between two runs on one machine when the cache is gone.

| decision | how it is made | remembered in | under `--deterministic` |
|---|---|---|---|
| the stacking degree | among the degrees whose memory plan fits the free memory: the only one, the cost model's pick when it is ahead by more than 5 percent, else a measurement at the first step; see [memory](https://sx.041.io/docs/devices.md#memory) | a measurement in `~/.cache/sexpgpu/choices.json` under the GPU, CUDA driver API version, loaded libraries, executable (its size and modification time), training graph, any generator's graphs and candidates; the cost model's timings in `~/.cache/sexpgpu/cost/<gpu>-<identity-hash>-<driver>.json` | 1, unstacked |
| a long F32 product's implementation | time ordinary, transposed, and legal split reductions; keep ordinary unless another is at least 1% faster; on several GPUs, ordinary unless pinned | `choices.json` under device, libraries, executable, dimensions, and operand layout; also recorded in the checkpoint's `products` map | ordinary pedantic GEMM, no measurement |
| a convolution's cuDNN engine | the fastest of its first eight deterministic plans, each timed once when the device first meets the role, shape and dtype | `~/.cache/sexpgpu/choices.json`, under the GPU, driver, cuDNN version, role, shape and dtype | the first deterministic plan cuDNN's heuristics offer |
| the fusion planner's recompute budget | the device's measured f32 flops over its bandwidth, which bounds how much arithmetic a fused reader recomputes instead of reading | `~/.cache/sexpgpu/cost/<gpu>-<identity-hash>-<driver>.json` | 9 flops a byte, the A100's |
| gradients between nodes | `SEXPGPU_NODE_GRADIENTS`, `bf16` by default; see [several nodes](https://sx.041.io/docs/devices.md#several-nodes) | nowhere; every node must set the same | `f32` |

Stacking changes the order in which F32 products of the microbatches sum;
two convolution engines sum in different orders. How far that goes on a
chaotic tier: tiny Adam at stacking 4 and at stacking 1 agree bit for bit
to step 10, differ by 2e-6 at step 20, and read an evaluation loss of 7.51
and 7.86 after 120 steps. The recompute budget changes which values are
recomputed, and the planner keeps every node's rounding when it
recomputes: tiny Adam at 9.00 and at 9.21 flops a byte agrees bit for bit.
It is recorded and pinned because its measurement moves, 9.21 to 10.99 on
one A100 between calibrations. bf16 between nodes rounds each device's
gradients and carries the rounding error into the next step; the carried
error is not in a checkpoint.

A [generator](https://sx.041.io/docs/generate.md) makes no choice of its own. Its graphs stack
with the training graph's degree, so the one stacking choice covers both;
it adds nothing to a checkpoint's identity beyond the file it is written
in; and its trip is an input like the step, so it repeats its bits
wherever the training graph does. A record pinned from a run without a
generator pins the stacking as any record does: taken when the run with
the generator offers that degree, `E-CHOICE-001` when the generator keeps
the run unstacked (it reads `:microbatch`, `:records` or `:stage-records`,
as a `:microbatch` salt does, or has `bf16` nodes) or the plan with the generator does not fit at that degree.

A run loses level 1 against an earlier run when:

- the cache was deleted, `HOME` is unset, or the cache cannot be written,
  and a measurement chooses otherwise; a pinned or `--deterministic` run
  reads no cache;
- another process holds device memory, so a different set of stacking
  degrees fits;
- an allocation fails during the run: the `memory: allocation failed`
  line says so, and a stacked run computes unstacked from that step. A
  pinned run stops with the allocation error instead;
- it resumes a run with bf16 between nodes: the resumed run starts
  without the carried error, one step's rounding away from the
  uninterrupted run. With `f32` between nodes, and on one node, a resume
  repeats the uninterrupted bits.

## Pinning a run to recorded choices

A run records every choice in the table above, the stacking in effect
and whether it ran under `--deterministic`, with the device facts that
level 2 needs to hold. They go into each checkpoint's `experiment.json`
under `choices` ([checkpoints](https://sx.041.io/docs/checkpoints.md#what-a-checkpoint-holds)),
into the metadata of the run's `lowering` annotation, and into the
`choices` line of the report; `sexpgpu checkpoint` prints them.

```json
"choices": {
  "device": "compute_80 with 108 multiprocessors, cuBLAS 130101, cuDNN 91300, NVRTC 13.0, patterns on",
  "strict": false,
  "stacking": 1,
  "recompute": 10.99...,
  "convolutions": {
    "Weight Spec { batch: 512, height: 8, width: 8, channels: 256, kh: 3, kw: 3, out: 512, stride: 1, padding: 1 } bf16": "engine 47 knobs 0=3 2=3 5=3 14=2",
    ...
  },
  "node_gradients": null
}
```

| flag | variable | the run's choices |
|---|---|---|
| none | none | its own; a resume takes its checkpoint's |
| `--choices <checkpoint>` | `SEXPGPU_CHOICES` | the ones that checkpoint recorded, a directory or `s3://bucket/prefix/step-<n>` |
| `--choices <file.json>` | `SEXPGPU_CHOICES` | the `choices` of a JSON file: an `experiment.json`, or the metadata of a `lowering` annotation |
| `--choices fresh` | `SEXPGPU_CHOICES=fresh` | its own, a resume too |
| `--deterministic` | `SEXPGPU_DETERMINISTIC=1` | none measured: the last column of the table above |

A pinned run measures nothing and remembers nothing. It takes each
recorded choice as it is, or stops before its first step with
`E-CHOICE-001` naming the choice it cannot take: a stacking whose plan does
not fit or is not offered, an engine cuDNN does not offer on this device
or a convolution the record has none for, or gradients between nodes in
another dtype than `SEXPGPU_NODE_GRADIENTS` says. It never falls back to
another choice. `--choices fresh` is the way past such an error, at level 3.

A resume pins itself to its checkpoint's choices, so a run preempted and
resumed on another machine keeps summing in its own order. A checkpoint
written before choices were recorded has none, and its resume chooses.
One written before product implementations were recorded has none of
them, and its resume takes ordinary GEMM, as its run did.
The `determinism` line of a pinned run says level 2 when the record's
device facts read as this machine's and level 3, naming both, when they
do not. A resume that chooses again, under `--choices fresh` or from a
checkpoint with no record, says level 3: its choices were made on another
machine than the steps before it. They often come out the same; a Track 3
soak's fresh resumes on A100s repeated the pinned soak's bits (see
`vision/decisions.md`, 2026-10-01), but nothing promises it.

`--deterministic` is slower: it gives up the stacking and the engines a
measurement would pick. It is for tests and for a comparison that must
hold without a record; the tests that assert bits between runs use it.
It refuses `--choices <record>` and `SEXPGPU_NODE_GRADIENTS=bf16` with
`E-CHOICE-002`.

## Offline evaluation

[`sexpgpu evaluate`](https://sx.041.io/docs/evaluation.md#sexpgpu-evaluate) runs a run's passes
from its checkpoints, often on another machine. A pass runs unstacked, so
the choices that move its bits are the recompute budget, convolution
engines, and CUDA F32 product implementations:

- On the kind of device the checkpoint's choices were made on, the
  recorded `device` with this one's facts at its start, it takes them and
  is at **level 2**: its `eval/*` values are the ones the run would have
  reported inline there. The CPU interpreter, which chooses nothing,
  repeats an inline CPU run's values bit for bit.
- On another kind, it chooses its own, measured and remembered as a run
  does, and is at **level 3**: the same distribution, not the same bits.
- `--choices` and `--deterministic` replace either, as for a run. A pinned
  convolution or eligible CUDA product whose shape has no recorded choice
  stops with `E-CHOICE-001`. A shape shared with training reuses its recorded
  implementation. `--choices fresh` chooses unrecorded shapes at level 3.

The `determinism` line of its lowering report and the `evaluated offline
on <device>: determinism <level>` annotation say which.

## What the device decides

Decisions that are fixed by the machine and its libraries, the same on
every run there, and different on another kind of machine. They are why
level 2 needs a machine that decides the lowering alike: the same compute
capability and multiprocessor count, the same cuBLAS, cuDNN and NVRTC,
and the same `SEXPGPU_PATTERNS`, which are the recorded `device`, with the
GPU and node counts appended when there are several.

| decision | decided by |
|---|---|
| cuBLAS's kernel for each product | cuBLAS, from the GPU and its version; an underfilled F32 batched product takes cuBLASLt's tiling, chosen by the multiprocessor count; the measured implementation of a long F32 product is a recorded choice listed above |
| the native cuDNN graphs: products, attention and convolution | used on `compute_80` with the pattern kernels on and cuDNN loaded, attention only with cuDNN 9.13.0; each takes the first plan cuDNN's heuristics offer, or for a convolution the engine above; elsewhere the ordinary lowering runs |
| the generated kernels' math functions | NVRTC and its libdevice for the device's compute capability; contraction into fused multiply-adds is off |
| the sum between GPUs | NCCL, from the topology and its version |

## What the run's configuration decides

These are part of the run, not the machine, and change its bits by
design: the GPU count and the node count (each rank sums its own
microbatches, NCCL sums across ranks), `SEXPGPU_PATTERNS`,
`SEXPGPU_OPTIMIZE`, and the diagnostics selection, whose sampled steps run
the diagnosed graph for their first microbatch.

## Random draws

A draw holds no state. Each element of `normal` or `uniform` is a pure
function of the node's stream key, which `defrun :seed` and the draw's
place in its graph decide, its salt when it has one, and the element's
index ([tensors](https://sx.041.io/docs/tensors.md#operations)). So a draw does not depend on the
machine, the device, the stacking or the order kernels run in: the
interpreter and CUDA give a uniform the same bits, and a normal the same
value up to the last ulp of the target's `ln`, `cos` and `sqrt`.

A salted draw is still such a function: the salt picks another stream and
nothing else. Nothing about a draw is in a checkpoint, and none needs to
be. A resume redraws exactly what the uninterrupted run drew, because the
salt is computed from values the resumed run has again: `:step` comes back
from the checkpoint, and `:microbatch` counts from zero in every step. That
holds while the resumed run keeps the run's microbatch count and its number
of GPUs: a salt that reads `:microbatch` draws for the global microbatch
index, and which data a microbatch index holds depends on both. A
salt read from a field changes between microbatches, so a graph with one is
not stacked; the separate calls it runs instead keep the bits.

What does change a draw: another `:seed`, another salt, another path for a
parameter's initializer, and in any other graph another draw added before
it, which renumbers the keys after it.

## Loaded weights

A `:load` parameter's starting value is a file, not the experiment file, so
the same file, selection and flags give the same bits only while that file
holds the same bytes. Each checkpoint's `experiment.json` pins it: the
parameter's `load` entry names the file read (the shard, for an index), the
tensor, and the file's size and version, its ETag on S3 or the SHA-256 of a
local file. A resume reads the file again and stops with `E-LOAD-001`,
showing both, when the tensor, size or version differ, rather than continue
from other weights; the same bytes at another path, a bundle's copy,
resume and are named in the `resumed from` annotation. A fresh run reads
whatever is at the path, so pin a source you mean to reproduce by never
writing over it: a new version under a new name. See
[checkpoints](https://sx.041.io/docs/checkpoints.md#what-a-checkpoint-holds).

Related: [devices](https://sx.041.io/docs/devices.md), [precision and numerics](https://sx.041.io/docs/numerics.md),
[checkpoints](https://sx.041.io/docs/checkpoints.md), [explain](https://sx.041.io/docs/explain.md#the-lowering-report).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
