# Diagnostic codes

Every diagnostic has a code, a primary span in a file you own, the expected
and actual facts, and a fix when one is known; see
[check](https://sx.041.io/docs/check.md#where-a-diagnostic-points). A wrong name gets the nearest
one that exists, a wrong count or keyword gets the signature, and a value of
the wrong kind is stated, type and text. `W-*` codes are warnings and
never stop a compilation. `sexpgpu check <file> --json` gives the same as
an array.

## Reading and evaluating

| code | what it means | usual fix |
|---|---|---|
| `E-READ-001` | the file ends inside a form or a string; the message names the form by its first line | add the closer; the note marks the first line inside it that starts at column one, where the closer most likely belongs |
| `E-READ-002` | a `)` or `]` with nothing open, or the wrong one for what is open | remove it, or close with the bracket the fix names; the note shows the open one |
| `E-READ-003` | a token that is not a number, keyword, string or symbol | `0.5` not `.5`; `:` plus a name for a keyword; a symbol is letters, digits and `-+*/<>=!?.&_%` |
| `E-EVAL-001` | the program called `(error ...)` | read the message |
| `E-EVAL-002` | an `(assert ...)` failed | read the message |
| `E-EVAL-003` | a name is not bound, or `setq` of one that is not; a knob used before its `defknobs` lands here. A name a macro wrote is looked up in the macro's module and the prelude, and the message names that module | the fix names the module that defines it, or the nearest name; for a name a macro wrote, define it in the macro's module or pass it as an argument |
| `E-EVAL-004` | calling something that is not a function, or a macro as a function | the span is the head and the message states its value; wrap a macro in a `lambda` |
| `E-EVAL-005` | a malformed special form, lambda list, or a binding of reserved syntax | the fix writes the form's shape; see [special forms](https://sx.041.io/docs/special-forms.md) |
| `E-EVAL-006` | an argument of the wrong type, a bad destructuring, a bad dtype keyword | the message states the value it found, type and text |
| `E-EVAL-007` | integer overflow or division by zero | |
| `E-EVAL-010` | a keyword argument that is not a keyword, unknown, or has no value | the fix names the nearest keyword and lists the accepted ones, or prints the signature; a note shows the definition |
| `E-EVAL-011` | the wrong number of arguments | the fix prints the signature, and a note shows the definition |
| `E-EVAL-030` | a file requires itself; the message spells the loop | move what both files need into a third file that both require |
| `E-EVAL-031` | a required file cannot be read; the message names the path it looked for | paths are relative to the file holding the `require` |
| `E-EVAL-032` | a require path is neither relative nor a package | `./name.sx`, or `sexpgpu/<module>` |
| `E-EVAL-033` | no such standard module | the fix lists the ones that exist |
| `E-EVAL-034` | a `require` names nothing or is in a local scope | list names, at module top level |
| `E-EVAL-035` | a `require` names what the file does not define | the fix is the nearest name it does |
| `E-EVAL-036` | a name bound twice in one file, by two imports (from one module or two) or an import and a definition | rename, or import once |
| `E-EVAL-037` | an import the file never uses: it writes the name nowhere outside its `require`s. A name written on a branch the selection skips is used, and so is a contract name imported into the run file | take the name out of the `require` |
| `E-EVAL-040` | recursion deeper than 256 calls; the message names the function and points at its recursive call | no tail calls; check the base case, or loop with `repeat`, `reduce` or `dotimes` |
| `E-EVAL-041` | compile-time work budget exceeded or invalid | raise `SEXPGPU_EVAL_LIMIT` for finite generators |

## Tensors and staging

| code | what it means | usual fix |
|---|---|---|
| `E-DIM-001` | shapes do not fit; the message names the axis or the dims that disagree. A `sum` or `max` that lists one axis twice, as `[1 -1]` on a rank-2 tensor, lands here too, as does a `table-scan` `:start` that is not a state of its table, and a `matrix-scan` whose `a` is not `[.., n, n]` on `b`'s `[.., n, m]` or whose `n` is above 16 | the note names both operands as you wrote them; the fix states the rule, or says to list each axis once |
| `E-DIM-002` | dtypes differ, or `:x` is not a dtype; a `table-scan` table or symbols that are not `:i32` | `(cast x :f32)`; dtypes are `:f32 :bf16 :i32 :i64 :bool` |
| `E-DIM-003` | a tensor was expected; the message states what was found | |
| `E-FLOW-001` | `if` on a tensor | select elementwise with `(where test then else)` |
| `E-FLOW-002` | a tensor from another graph; the message names its first binding | recompute it in this graph instead of keeping it in a variable |
| `E-FLOW-003` | a tensor operation with no graph being traced | tensors exist in `prepare`, a model, `objective`, `evaluate`, an optimizer body, an initializer or `curriculum` |
| `E-GRAD-001` | `(gradient y x)` where `y` is not one float value; the message gives its shape and dtype | reduce `y` with `sum` or `mean` first |
| `E-GRAD-002` | `(gradient y x)` where `x` is an integer or boolean tensor | take it with respect to a float tensor `y` is computed from |
| `E-GRAD-003` | `(gradient y x)` where `y` is not computed from `x`, so the gradient would be zeros | differentiate with respect to a tensor `y` reads, or write `(zeros-like x)` |
| `E-GRAD-004` | `(gradient y x)` where `y` reads the gradient a `tap-gradient` lambda is given, which is filled in after the training backward | take `gradient` of a value computed without it |

## Models and optimizers

| code | what it means | usual fix |
|---|---|---|
| `E-PARAM-002` | `defparam` outside a model | parameters live inside `defmodel` |
| `E-PARAM-003` | an initializer reads another value | an initializer is closed: shapes, constants, random nodes |
| `E-PARAM-004` | a model body does not end in a function; the message shows the form it ends in | return `(lambda (x) ...)` last |
| `E-PARAM-005` | `named` on something that is not a model or a parameter | |
| `E-PARAM-006` | an unknown parameter option, or `:numerics` other than `:high` | the fix names the nearest; the options are `:tags`, `:numerics`, `:trainable` and `:load` |
| `E-PARAM-007` | `:trainable` that is not a boolean, or `:load` that is not a file and a tensor name; a `with-loaded` header that is not a file, a function and an optional `:trainable` boolean, or its function returning anything but a string for a path, which the message names with the value | `:trainable false`, `:load ["model.safetensors" "name"]`, a name function that returns a string |
| `E-PARAM-008` | a model under two parameter roots: bound to two top-level names, or reachable from two roots; the note shows the other binding | bind each model to one top-level name |
| `E-PARAM-009` | a graph reads a parameter of a model not bound by itself at the top level: built in a `let`, a function or a top-level list | bind it with `defvar` at the top of the run file, or build it inside the `defmodel` that uses it |
| `W-PARAM-001` | a model is never bound to a name | bind it with `let`, or `named` |
| `E-OPT-001` | a selector matches no parameter | the fix names the nearest parameter path and lists them; or widen the glob |
| `E-OPT-002` | a parameter is in more than one group | narrow one selector with `select-not` |
| `E-OPT-003` | a misplaced or duplicate state, or a state initializer that is not a zero constant | start it at zero: `zeros-like`, `zeros` or a zero `full` |
| `E-OPT-004` | an optimizer body does not end in `optimizer-update`, or drops a state; the message lists declared and returned states | return every declared state exactly once |
| `E-OPT-005` | a malformed `select`, `group`, `:groups` or `:optimizer`, or an inner optimizer with `:groups`; the message states what was given | the fix prints the shape |
| `E-OBJ-001` | `objective` does not return one float scalar | reduce the loss; report the parts with `metric` |

## The contract

| code | what it means | usual fix |
|---|---|---|
| `E-CONTRACT-001` | a required top-level name is missing | the fix writes the form that defines it, `(defrun :steps 300 :microbatches 4)`; see [the run file](https://sx.041.io/docs/run-file.md) |
| `E-CONTRACT-002` | a contract name is the wrong kind of value; the span is where the file defines or imports it | the fix writes the form that defines it |
| `E-CONTRACT-003` | an unknown keyword in a configuration form or a builtin's options, a bad `:precision` or `:evaluate`, a pass's `:prepare` or `:evaluate` that is not a function, or a `manifest` `:subsets` that is not a non-empty vector of strings | the fix names the nearest keyword and lists the accepted ones, or writes `:subsets ["context:512"]` |
| `E-CONTRACT-004` | `prepare` did not return `:inputs` and `:targets`; the message lists the keys it did | `(list :inputs ... :targets ...)`; the fix names a near miss |
| `E-CONTRACT-005` | the batch has no such field | the fix lists what the loader declares |
| `E-CONTRACT-010` | `counter` outside `prepare` | count where records are counted |
| `E-CONTRACT-011` | `metric` in an optimizer body | a `diagnostic`, or compute it in `objective` or `evaluate` |
| `E-CONTRACT-012` | a metric's metadata is over what Metrics accepts, runtime keys counted | fewer or shorter keys |
| `E-CONTRACT-013` | two `defeval` passes share a name | |
| `E-CONTRACT-014` | `evaluate` or `curriculum` with no `defeval` pass | declare a pass |
| `E-CONTRACT-015` | a `diagnostic` without `:reduce`, with an unknown one, or two reducers for one name | one of `:mean :min :max :sum` |
| `E-CONTRACT-016` | a diagnostics selection that names nothing the file defines | the fix lists the names |
| `E-CONTRACT-017` | a `metric` inside a `tap-gradient` lambda | report the gradient with `diagnostic` |
| `E-CONTRACT-018` | `defrun :evaluate :offline` (or `run --evaluate offline`) whose `:checkpoint-every` does not divide a pass's `:every`, or none while a pass has one: `sexpgpu evaluate` would miss that pass's steps | a `:checkpoint-every` that divides every `:every`, or evaluate `:inline`; see [evaluation](https://sx.041.io/docs/evaluation.md#when-a-pass-runs) |

## Generation

| code | what it means | usual fix |
|---|---|---|
| `E-GEN-001` | a generator's `:init`, `:step` or `:finish` returned something other than a plist of tensors; the span is the lambda, the message shows what it returned | `(list :sequence s :cache c)`, a keyword before each tensor |
| `E-GEN-002` | `:step` returned other keys, shapes or dtypes than `:init` made; the message shows both | the state keeps its keys, shapes and dtypes on every trip |
| `E-GEN-003` | the generator makes a name the training loader or a pass's loader reads as a field; the span is the lambda that returned it. A pass whose loader has other fields or another batch size than the training loader's does not run the generator, so its `prepare` reading a generated name is `E-CONTRACT-005` | rename it where the generator returns it; give the pass the training loader's fields and batch size |
| `E-GEN-004` | a `metric` inside a generator, which would never be reported. A `diagnostic` there, as in a model the generator runs, is never selected and is dropped; a diagnostics selection naming only such a one is `E-CONTRACT-016` | return the value in the state and report it from `objective` |
| `E-GEN-005` | `defgenerate` without `:trips`, `:init` or `:step`, or `:trips` below 1; the message shows the value | `(defgenerate :trips 15 :init start :step next)` |

## Knobs

| code | what it means | usual fix |
|---|---|---|
| `E-KNOB-001` | a knob is declared twice | the note shows the first; delete one |
| `E-KNOB-002` | a knob the file declares and never reads; a variant, a sweep or `--set` setting it is not a read | read it, or delete it |
| `E-KNOB-003` | a variant or sweep sets something that is not a knob | the fix names the nearest and lists the knobs |
| `E-KNOB-004` | `--variant` or a sweep names a variant the file does not declare; the span is the nearest declaration | the fix names the nearest and lists the variants |
| `E-KNOB-005` | `--set` names a knob the file does not declare; the span is the nearest declaration | the fix names the nearest and lists the knobs |
| `E-KNOB-006` | a variant declared twice, or after `defknobs` | the note shows the first or the knobs; move every `defvariant` above `defknobs` |
| `E-KNOB-007` | a sweep axis lists one value twice; the two points would be one run under one slug. The message says which values | drop the repeat |

## Run time and tools

| code | what it means | usual fix |
|---|---|---|
| `E-DP-001` .. `E-DP-007` | several GPUs or nodes: configuration, rendezvous and resume | see [devices](https://sx.041.io/docs/devices.md#several-gpus) |
| `E-CHOICE-001` | a pinned run cannot take a recorded choice here: a stacking that does not fit or is not offered, a cuDNN engine not offered on this device or a convolution the record has none for, another dtype between nodes; the message names it | a device like the recorded one, `SEXPGPU_NODE_GRADIENTS` as recorded, or `--choices fresh`; see [determinism](https://sx.041.io/docs/determinism.md#pinning-a-run-to-recorded-choices) |
| `E-CHOICE-002` | `--choices` names no checkpoint or file with recorded choices, or `--deterministic` was given with `--choices <record>` or `SEXPGPU_NODE_GRADIENTS=bf16` | a checkpoint written by this version, or one of the two |
| `E-LOAD-001` | a resume whose `:load` file is not the one the checkpoint's run read: another tensor, size or version (ETag, or SHA-256 of a local file); the message shows both | restore the file, or start the run again; see [checkpoints](https://sx.041.io/docs/checkpoints.md#resume) |
| `E-MEM-001` | the smallest memory plan does not fit on the device; the message names the number and the largest part | fewer rows per microbatch (raise `:microbatches`), or a device with more memory; across nodes, with f32 gradients and bf16 transfer the collective buffers cost 10 bytes per parameter per GPU and `SEXPGPU_NODE_GRADIENTS=f32` saves 6 of them; see [devices](https://sx.041.io/docs/devices.md#memory) |
| `E-MEM-002` | a required buffer exceeds Metal's advertised per-buffer limit, or allocation fails; the message names the bytes and the tensor | a smaller shape or more `:microbatches`; Metal has no memory plan, so this is its only check; see [devices](https://sx.041.io/docs/devices.md#the-metal-device) |
| `E-MEM-003` | a CUDA training microbatch still could not allocate after the device's caches were released: its retry failed, or it was the third microbatch in a row to fail. The run needs more memory than the plan said. The message names the step, the microbatch, the planned memory and what is free now | fewer rows per microbatch (raise `:microbatches`), a smaller model, or a device with more memory; see [devices](https://sx.041.io/docs/devices.md#memory) |
| `E-FMT-001` | the formatter refused its own output | a compiler bug; the file was left as it was |
| `E-INTERNAL-001` | the lowered document failed validation | a compiler bug, never the experiment's fault |

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
