S-exp GPU
All pages
Docs · Models and optimizersMarkdown

Determinism

What a run may rely on about its bits. Every kernel SexpGPU runs sums in a fixed order: there are no floating-point atomics, the cuDNN engines it accepts are the deterministic ones, and cuBLAS keeps its default of no atomics. What can still move the bits is a decision made from the machine rather than from the file. This page names each of them, where it is remembered, and the level of reproduction a run gets.

Which flag, when

you wantrun with
the fastest run, on whatever card you getnothing: the run measures its choices, remembers them in ~/.cache/sexpgpu, and repeats its own bits on that machine (level 1)
a spot job that may resume on another card--choices fresh, or SEXPGPU_CHOICES=fresh in the job's environment: a resume chooses again where its checkpoint's choices do not fit or are not offered, and continues at level 3 rather than stopping
a resume that keeps the exact bits of the run it continuesnothing, on the same kind of card: a resume pins itself to its checkpoint's choices (level 2), and stops with E-CHOICE-001 rather than sum in another order if the card cannot take them
the same bits as a colleague's run, on your card--choices <their checkpoint or experiment.json>
bits that hold with no record at all, for a test or a paper's comparison--deterministic: a few percent slower, nothing measured

The three levels

  1. One device, same bits. The same file, selection and flags, run twice on one machine with the same binary, GPUs, driver, libraries and ~/.cache/sexpgpu, give every loss, metric, parameter and optimizer state the same bits. Every run has this level. The cache is what makes it true: a decision measured once is remembered and taken again, instead of measured again and possibly taken otherwise.
  2. Any device, same bits, given the choices. A run pinned to another run's recorded choices reproduces its bits on another machine, provided that machine can execute every choice and decides the rest of the lowering alike (below). A run under --deterministic has this level against every other --deterministic run of the file, without any record: its choices are fixed rules.
  3. Same distribution. Without the same choices another machine may decide differently, and its sums run in another order. The loss curve stays inside the ladder's envelope for its tier, not bit for bit.

The CPU interpreter decides nothing from the machine; a run on it is at level 1 on any machine with the same platform math library.

Metal makes no measured choices. It uses stacking 1, zero recomputation, and the IR composition for convolutions. Its checkpoint identity includes the Apple GPU name, MPS F32 schedule with K at most 1024, batched leading sums with at least 128 matrices and 65536 elements per product, independent matrix batching, and macOS build. Independent batching inserts the literal independent-batches between M*N>=65536 and macOS in choices.device. Lowering annotations repeat that identity in metadata.device and metadata.choices.device. The OS supplies the Metal compiler, driver, and MPS implementation. Repeated graphs and checkpoint resume reproduce bits on the same configuration. CPU and CUDA comparisons use numerical tolerances; Metal's BF16 storage follows the CPU interpreter's cast boundaries. Eligible dense F32 matrix products use MPS with reduced precision disabled; their floating-point association can differ from the tiled shaders and the interpreter. Batched and per-matrix MPS calls can also produce different product bits. Equality measured on one device does not guarantee equality on another GPU or macOS build, where MPS can use a different schedule. SIMD-group products, parallel reductions, and segmented scans use fixed schedules that can round differently from the interpreter's sequential folds. Batch sums of tiled products retain each product's rounding and the sum's order, as when the products are materialized separately.

The lowering report of a CUDA run ends with a determinism line that names the run's level and why, and a choices line with what it chose:

  determinism         1: this machine repeats these bits with this binary and ~/.cache/sexpgpu; another is 3, or 2 pinned to these choices with --choices
  choices             stacking 1, recompute 10.99 flops a byte, 17 convolution engines; on compute_80 with 108 multiprocessors, cuBLAS 130101, cuDNN 91300, NVRTC 13.0, patterns on

What a run chooses

Decisions made by measuring, or from what the machine holds at the time. They are the ones that can differ between two machines with the same card, and between two runs on one machine when the cache is gone.

decisionhow it is maderemembered inunder --deterministic
the stacking degreeamong the degrees whose memory plan fits the free memory: the only one, the cost model's pick when it is ahead by more than 5 percent, else a measurement at the first step; see memorya measurement in ~/.cache/sexpgpu/choices.json under the GPU, CUDA driver API version, loaded libraries, executable (its size and modification time), training graph, any generator's graphs and candidates; the cost model's timings in ~/.cache/sexpgpu/cost/<gpu>-<identity-hash>-<driver>.json1, unstacked
a long F32 product's implementationtime ordinary, transposed, and legal split reductions; keep ordinary unless another is at least 1% faster; on several GPUs, ordinary unless pinnedchoices.json under device, libraries, executable, dimensions, and operand layout; also recorded in the checkpoint's products mapordinary pedantic GEMM, no measurement
a convolution's cuDNN enginethe fastest of its first eight deterministic plans, each timed once when the device first meets the role, shape and dtype~/.cache/sexpgpu/choices.json, under the GPU, driver, cuDNN version, role, shape and dtypethe first deterministic plan cuDNN's heuristics offer
the fusion planner's recompute budgetthe device's measured f32 flops over its bandwidth, which bounds how much arithmetic a fused reader recomputes instead of reading~/.cache/sexpgpu/cost/<gpu>-<identity-hash>-<driver>.json9 flops a byte, the A100's
gradients between nodesSEXPGPU_NODE_GRADIENTS, bf16 by default; see several nodesnowhere; every node must set the samef32

Stacking changes the order in which F32 products of the microbatches sum; two convolution engines sum in different orders. How far that goes on a chaotic tier: tiny Adam at stacking 4 and at stacking 1 agree bit for bit to step 10, differ by 2e-6 at step 20, and read an evaluation loss of 7.51 and 7.86 after 120 steps. The recompute budget changes which values are recomputed, and the planner keeps every node's rounding when it recomputes: tiny Adam at 9.00 and at 9.21 flops a byte agrees bit for bit. It is recorded and pinned because its measurement moves, 9.21 to 10.99 on one A100 between calibrations. bf16 between nodes rounds each device's gradients and carries the rounding error into the next step; the carried error is not in a checkpoint.

A generator makes no choice of its own. Its graphs stack with the training graph's degree, so the one stacking choice covers both; it adds nothing to a checkpoint's identity beyond the file it is written in; and its trip is an input like the step, so it repeats its bits wherever the training graph does. A record pinned from a run without a generator pins the stacking as any record does: taken when the run with the generator offers that degree, E-CHOICE-001 when the generator keeps the run unstacked (it reads :microbatch, :records or :stage-records, as a :microbatch salt does, or has bf16 nodes) or the plan with the generator does not fit at that degree.

A run loses level 1 against an earlier run when:

  • the cache was deleted, HOME is unset, or the cache cannot be written, and a measurement chooses otherwise; a pinned or --deterministic run reads no cache;
  • another process holds device memory, so a different set of stacking degrees fits;
  • an allocation fails during the run: the memory: allocation failed line says so, and a stacked run computes unstacked from that step. A pinned run stops with the allocation error instead;
  • it resumes a run with bf16 between nodes: the resumed run starts without the carried error, one step's rounding away from the uninterrupted run. With f32 between nodes, and on one node, a resume repeats the uninterrupted bits.

Pinning a run to recorded choices

A run records every choice in the table above, the stacking in effect and whether it ran under --deterministic, with the device facts that level 2 needs to hold. They go into each checkpoint's experiment.json under choices (checkpoints), into the metadata of the run's lowering annotation, and into the choices line of the report; sexpgpu checkpoint prints them.

"choices": {
  "device": "compute_80 with 108 multiprocessors, cuBLAS 130101, cuDNN 91300, NVRTC 13.0, patterns on",
  "strict": false,
  "stacking": 1,
  "recompute": 10.99...,
  "convolutions": {
    "Weight Spec { batch: 512, height: 8, width: 8, channels: 256, kh: 3, kw: 3, out: 512, stride: 1, padding: 1 } bf16": "engine 47 knobs 0=3 2=3 5=3 14=2",
    ...
  },
  "node_gradients": null
}
flagvariablethe run's choices
nonenoneits own; a resume takes its checkpoint's
--choices <checkpoint>SEXPGPU_CHOICESthe ones that checkpoint recorded, a directory or s3://bucket/prefix/step-<n>
--choices <file.json>SEXPGPU_CHOICESthe choices of a JSON file: an experiment.json, or the metadata of a lowering annotation
--choices freshSEXPGPU_CHOICES=freshits own, a resume too
--deterministicSEXPGPU_DETERMINISTIC=1none measured: the last column of the table above

A pinned run measures nothing and remembers nothing. It takes each recorded choice as it is, or stops before its first step with E-CHOICE-001 naming the choice it cannot take: a stacking whose plan does not fit or is not offered, an engine cuDNN does not offer on this device or a convolution the record has none for, or gradients between nodes in another dtype than SEXPGPU_NODE_GRADIENTS says. It never falls back to another choice. --choices fresh is the way past such an error, at level 3.

A resume pins itself to its checkpoint's choices, so a run preempted and resumed on another machine keeps summing in its own order. A checkpoint written before choices were recorded has none, and its resume chooses. One written before product implementations were recorded has none of them, and its resume takes ordinary GEMM, as its run did. The determinism line of a pinned run says level 2 when the record's device facts read as this machine's and level 3, naming both, when they do not. A resume that chooses again, under --choices fresh or from a checkpoint with no record, says level 3: its choices were made on another machine than the steps before it. They often come out the same; a Track 3 soak's fresh resumes on A100s repeated the pinned soak's bits (see vision/decisions.md, 2026-10-01), but nothing promises it.

--deterministic is slower: it gives up the stacking and the engines a measurement would pick. It is for tests and for a comparison that must hold without a record; the tests that assert bits between runs use it. It refuses --choices <record> and SEXPGPU_NODE_GRADIENTS=bf16 with E-CHOICE-002.

Offline evaluation

sexpgpu evaluate runs a run's passes from its checkpoints, often on another machine. A pass runs unstacked, so the choices that move its bits are the recompute budget, convolution engines, and CUDA F32 product implementations:

  • On the kind of device the checkpoint's choices were made on, the recorded device with this one's facts at its start, it takes them and is at level 2: its eval/* values are the ones the run would have reported inline there. The CPU interpreter, which chooses nothing, repeats an inline CPU run's values bit for bit.
  • On another kind, it chooses its own, measured and remembered as a run does, and is at level 3: the same distribution, not the same bits.
  • --choices and --deterministic replace either, as for a run. A pinned convolution or eligible CUDA product whose shape has no recorded choice stops with E-CHOICE-001. A shape shared with training reuses its recorded implementation. --choices fresh chooses unrecorded shapes at level 3.

The determinism line of its lowering report and the evaluated offline on <device>: determinism <level> annotation say which.

What the device decides

Decisions that are fixed by the machine and its libraries, the same on every run there, and different on another kind of machine. They are why level 2 needs a machine that decides the lowering alike: the same compute capability and multiprocessor count, the same cuBLAS, cuDNN and NVRTC, and the same SEXPGPU_PATTERNS, which are the recorded device, with the GPU and node counts appended when there are several.

decisiondecided by
cuBLAS's kernel for each productcuBLAS, from the GPU and its version; an underfilled F32 batched product takes cuBLASLt's tiling, chosen by the multiprocessor count; the measured implementation of a long F32 product is a recorded choice listed above
the native cuDNN graphs: products, attention and convolutionused on compute_80 with the pattern kernels on and cuDNN loaded, attention only with cuDNN 9.13.0; each takes the first plan cuDNN's heuristics offer, or for a convolution the engine above; elsewhere the ordinary lowering runs
the generated kernels' math functionsNVRTC and its libdevice for the device's compute capability; contraction into fused multiply-adds is off
the sum between GPUsNCCL, from the topology and its version

What the run's configuration decides

These are part of the run, not the machine, and change its bits by design: the GPU count and the node count (each rank sums its own microbatches, NCCL sums across ranks), SEXPGPU_PATTERNS, SEXPGPU_OPTIMIZE, and the diagnostics selection, whose sampled steps run the diagnosed graph for their first microbatch.

Random draws

A draw holds no state. Each element of normal or uniform is a pure function of the node's stream key, which defrun :seed and the draw's place in its graph decide, its salt when it has one, and the element's index (tensors). So a draw does not depend on the machine, the device, the stacking or the order kernels run in: the interpreter and CUDA give a uniform the same bits, and a normal the same value up to the last ulp of the target's ln, cos and sqrt.

A salted draw is still such a function: the salt picks another stream and nothing else. Nothing about a draw is in a checkpoint, and none needs to be. A resume redraws exactly what the uninterrupted run drew, because the salt is computed from values the resumed run has again: :step comes back from the checkpoint, and :microbatch counts from zero in every step. That holds while the resumed run keeps the run's microbatch count and its number of GPUs: a salt that reads :microbatch draws for the global microbatch index, and which data a microbatch index holds depends on both. A salt read from a field changes between microbatches, so a graph with one is not stacked; the separate calls it runs instead keep the bits.

What does change a draw: another :seed, another salt, another path for a parameter's initializer, and in any other graph another draw added before it, which renumbers the keys after it.

Loaded weights

A :load parameter's starting value is a file, not the experiment file, so the same file, selection and flags give the same bits only while that file holds the same bytes. Each checkpoint's experiment.json pins it: the parameter's load entry names the file read (the shard, for an index), the tensor, and the file's size and version, its ETag on S3 or the SHA-256 of a local file. A resume reads the file again and stops with E-LOAD-001, showing both, when the tensor, size or version differ, rather than continue from other weights; the same bytes at another path, a bundle's copy, resume and are named in the resumed from annotation. A fresh run reads whatever is at the path, so pin a source you mean to reproduce by never writing over it: a new version under a new name. See checkpoints.

Related: devices, precision and numerics, checkpoints, explain.