# Evaluation passes

A pass is a named loader and a cadence. Every pass runs `eval-prepare` (or
`prepare`), the model and [`evaluate`](https://sx.041.io/docs/objective.md#evaluate) over its own
loader, or the `:prepare` and `:evaluate` it names. A file declares as many
passes as it needs.

A pass is a pass; where it runs is a separate choice. Inline, the
training loop runs it between steps. Offline, the run only writes
checkpoints and `sexpgpu evaluate` runs the same pass from them, in
another process and perhaps on another machine and another kind of
device, reporting the same series at the same steps.

```lisp
(defeval eval
  :loader (loader :sources [(files ["data/val.parquet"])]
                  :fields [(field :tokens :from "input_ids" :dtype :i32 :shape [1025])]
                  :batch-size 4)
  :every 25)

(defeval probe :loader val-loader :every 5 :batches 16)
```

| keyword | default | meaning |
|---|---|---|
| `:loader` | required | the records the pass reads; see [loaders](https://sx.041.io/docs/data.md) |
| `:every` | none | steps between runs; without it the pass runs only after the last step |
| `:batches` | the whole loader | read this many batches from the start; without it the loader must be finite |
| `:prepare` | `eval-prepare`, then `prepare` | this pass's own prepare, for a loader whose records the file's are not written for |
| `:evaluate` | `evaluate` | this pass's own evaluate; the file needs no `evaluate` when every pass has one |

## When a pass runs

`defrun :evaluate` says where every pass of the file runs:
`:inline`, the default, or `:offline`. `sexpgpu run --evaluate
inline|offline` replaces it for one session, recorded as knob
`evaluate`, source `override`.

Inline:

- At step 0, every `:every` steps, and once more after the last step.
- Passes due at the same step run in the order the file declares them.
- A pass reopens its loader every time, so a limited pass reads the same
  records each time.
- A pass whose loader has the training loader's fields and batch size runs
  the file's [generator](https://sx.041.io/docs/generate.md) on each batch, unstacked, and reads
  what it makes as `prepare` does in training.
- `--eval-every N` replaces every pass's `:every` for one session; `0`
  leaves only the pass after the last step, so a curriculum never runs. See
  [run](https://sx.041.io/docs/run.md#flags).

Offline:

- The run prepares no pass graph and runs none, not even after the last
  step; its memory plan holds none, and on a GPU its
  [lowering report](https://sx.041.io/docs/explain.md#the-lowering-report) says
  `evaluation  offline: eval, probe run from the checkpoints with sexpgpu
  evaluate`. A curriculum never runs.
- `:every` keeps its meaning, the interval the pass wants, and is honoured
  against the checkpoints that exist: `evaluate --follow` runs a pass on a
  checkpoint whose step is a multiple of its `:every`. A pass without
  `:every` runs on every checkpoint, and every pass runs on the checkpoint
  of the last step. There is no checkpoint of step 0.
- So every `:every` must be a multiple of `:checkpoint-every`: `check`
  refuses a file whose checkpoints do not land on a pass's `:every`
  (`:every 5` under `:checkpoint-every 2` would be evaluated only every
  10), or that has none while a pass has an `:every` (`E-CONTRACT-018`).
  `run --evaluate offline` refuses the same, and refuses to start without
  a checkpoint location.

```lisp
(defeval eval :loader val-loader :every 250)
(defeval probe :loader val-loader :every 50 :batches 16)

(defrun :steps 3000 :checkpoint-every 50 :evaluate :offline)
```

## sexpgpu evaluate

```bash
sexpgpu evaluate my-run.sx --checkpoint s3://my-bucket/runs/my-run/step-00000300
sexpgpu evaluate my-run.sx --checkpoint s3://my-bucket/runs/my-run --follow --pass probe
```

It compiles the file under the checkpoint's selection, its variant and
every knob with the value `experiment.json` records, and the run's
`--steps`, `--eval-every`, `--evaluate` and `--diagnostics` overrides, so
the passes are the ones the run would have run and an unedited file shows
no change; `--variant` and `--set` are not taken.
It verifies the file against the checkpoint as a
[resume](https://sx.041.io/docs/checkpoints.md#resume) does: a parameter that is new, gone, of
another shape or dtype, or trains where it was frozen is refused, naming
each one, and so is a `:load` file whose bytes changed (`E-LOAD-001`);
edited sources are allowed and printed under the checkpoint, one line
each. Then it builds the pass graphs alone, and the generator's when a
pass runs it; a `with-loaded` teacher is read from the files the
checkpoint pinned. The trained parameters come from `params.safetensors`;
the optimizer states and the training loader are never read. Each pass
reports into the run's own Metrics experiment, named by the slug
`experiment.json` records, at the checkpoint's step, with the series names
and metadata it has inline; a checkpoint without a slug is refused. The
experiment keeps describing the training run: its `open` carries the
checkpoint's manifest, every pass whatever `--pass` picked, and the
device kind and count the checkpoint's choices record. See [metrics](https://sx.041.io/docs/metrics.md#offline-evaluation).

| flag | meaning |
|---|---|
| `--checkpoint <dir>\|s3://bucket/prefix/step-<n>` | evaluate that checkpoint |
| `--checkpoint <dir>\|s3://bucket/prefix` | a directory of `step-<n>` checkpoints: the newest complete one, or with `--follow` each as it appears |
| `--checkpoint latest` | the same under `--checkpoint-dir` or `SEXPGPU_CHECKPOINT_DIR` |
| `--pass <name>` | only this pass; repeatable. Every pass by default; an undeclared one is an error |
| `--follow` | watch the directory: evaluate each new checkpoint once, in step order, until the run's last |
| `--poll <seconds>` | how often `--follow` lists the directory; `60` |
| `--device`, `--metrics`, `--choices`, `--deterministic`, `--checkpoint-dir` | as for [run](https://sx.041.io/docs/run.md#flags); `SEXPGPU_DEVICES` picks the GPU, its first ordinal, so an evaluation beside a training run can keep off the run's |

Without `--follow`, every selected pass runs on the one checkpoint,
whatever its `:every`. A machine with no GPU evaluates small passes with
`--device cpu`.

**Choices.** The checkpoint's recorded choices are taken when this is the
kind of device they were made on, and the evaluation is at
[level 2](https://sx.041.io/docs/determinism.md#offline-evaluation): the CPU interpreter
repeats an inline CPU run's `eval/*` values bit for bit. On another kind
of device it chooses its own and is at level 3. `--choices` and
`--deterministic` replace either. The lowering report and the first
annotation say which.

**Memory.** On a GPU the memory plan holds the parameters, the loaded
teacher among them, the generator's phases and each pass's graph; no
optimizer state, gradient or training graph. A pass that does not fit
stops before it starts, `E-MEM-001`, as a run does; see
[devices](https://sx.041.io/docs/devices.md#memory).

**Following.** `--follow` remembers each pass's evaluated steps in
`evaluated.json` beside the checkpoints (see
[checkpoints](https://sx.041.io/docs/checkpoints.md#what-evaluate-reads)), written after each
pass once its numbers have reached the metrics target, so a follower
restarted on any machine repeats no step it did and loses none it
marked. Numbers that do not arrive within 60 seconds stop the follower
with an error, the step undone. It
stops after the checkpoint whose `experiment.json` says `"finished":
true`, which a run writes in the checkpoint of its last step. A follower started before
the first checkpoint waits for it.

**What it never does.** It writes no checkpoint, no status file, and
never starts or finishes the experiment: the run's state in Metrics is the
training run's. One evaluation can run beside the training run, or long
after it.

**Stopping.** `SIGTERM`, `SIGINT` or `SIGHUP` stops a pass at its next
batch; that pass reports nothing, and in `--follow` its step stays undone
for a restart. A follower waiting for a checkpoint stops at once. The
process prints `sexpgpu evaluate: interrupted by <signal>` and exits
`128 + n`.

**Output.** On standard error: where the numbers go, the lowering report
on a GPU, `evaluate: on <device>, determinism <level>`, one `evaluate:
<checkpoint>` line per checkpoint with any edited source under it, one
[progress line](https://sx.041.io/docs/run.md#output) per pass, and a `done:` block with each
pass's numbers.

| code | when |
|---|---|
| `0` | every pass asked for ran |
| `1` | refused or failed: diagnostics, an identity that does not match, no checkpoint, an undeclared `--pass`, a pass that did not fit |
| `2` | the command line was wrong: no file, no `--checkpoint`, a `--poll` that is not a positive number |
| `128 + n` | stopped by signal `n` |

## The pass named eval

The pass named `eval` is the one the summary, the status file, the progress
line and `sweep`'s table report; a file without one reports its first pass.
Every observation from a pass carries `pass` metadata, so `eval/loss` from
`eval` and from `probe` are two series. See
[metrics](https://sx.041.io/docs/metrics.md#runtime-metadata).

## Errors

| code | when |
|---|---|
| `E-CONTRACT-013` | two passes share a name |
| `E-CONTRACT-014` | an `evaluate` or a `curriculum` in a file with no pass, so nothing would run it |
| `E-CONTRACT-003` | a pass's `:prepare` or `:evaluate` that is not a function, a `defrun :evaluate` other than `:inline` or `:offline` |
| `E-CONTRACT-018` | an offline run whose checkpoints do not land on a pass's `:every` |

Related: [curriculum](https://sx.041.io/docs/curriculum.md), [the status file](https://sx.041.io/docs/status.md),
[checkpoints](https://sx.041.io/docs/checkpoints.md), [determinism](https://sx.041.io/docs/determinism.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
