All pages
Evaluation passes
A pass is a named loader and a cadence. Every pass runs eval-prepare (or
prepare), the model and evaluate over its own
loader, or the :prepare and :evaluate it names. A file declares as many
passes as it needs.
A pass is a pass; where it runs is a separate choice. Inline, the
training loop runs it between steps. Offline, the run only writes
checkpoints and sexpgpu evaluate runs the same pass from them, in
another process and perhaps on another machine and another kind of
device, reporting the same series at the same steps.
(defeval eval
:loader (loader :sources [(files ["data/val.parquet"])]
:fields [(field :tokens :from "input_ids" :dtype :i32 :shape [1025])]
:batch-size 4)
:every 25)
(defeval probe :loader val-loader :every 5 :batches 16)
| keyword | default | meaning |
|---|---|---|
:loader | required | the records the pass reads; see loaders |
:every | none | steps between runs; without it the pass runs only after the last step |
:batches | the whole loader | read this many batches from the start; without it the loader must be finite |
:prepare | eval-prepare, then prepare | this pass's own prepare, for a loader whose records the file's are not written for |
:evaluate | evaluate | this pass's own evaluate; the file needs no evaluate when every pass has one |
When a pass runs
defrun :evaluate says where every pass of the file runs:
:inline, the default, or :offline. sexpgpu run --evaluate inline|offline replaces it for one session, recorded as knob
evaluate, source override.
Inline:
- At step 0, every
:everysteps, and once more after the last step. - Passes due at the same step run in the order the file declares them.
- A pass reopens its loader every time, so a limited pass reads the same records each time.
- A pass whose loader has the training loader's fields and batch size runs
the file's generator on each batch, unstacked, and reads
what it makes as
preparedoes in training. --eval-every Nreplaces every pass's:everyfor one session;0leaves only the pass after the last step, so a curriculum never runs. See run.
Offline:
- The run prepares no pass graph and runs none, not even after the last
step; its memory plan holds none, and on a GPU its
lowering report says
evaluation offline: eval, probe run from the checkpoints with sexpgpu evaluate. A curriculum never runs. :everykeeps its meaning, the interval the pass wants, and is honoured against the checkpoints that exist:evaluate --followruns a pass on a checkpoint whose step is a multiple of its:every. A pass without:everyruns on every checkpoint, and every pass runs on the checkpoint of the last step. There is no checkpoint of step 0.- So every
:everymust be a multiple of:checkpoint-every:checkrefuses a file whose checkpoints do not land on a pass's:every(:every 5under:checkpoint-every 2would be evaluated only every 10), or that has none while a pass has an:every(E-CONTRACT-018).run --evaluate offlinerefuses the same, and refuses to start without a checkpoint location.
(defeval eval :loader val-loader :every 250)
(defeval probe :loader val-loader :every 50 :batches 16)
(defrun :steps 3000 :checkpoint-every 50 :evaluate :offline)
sexpgpu evaluate
sexpgpu evaluate my-run.sx --checkpoint s3://my-bucket/runs/my-run/step-00000300
sexpgpu evaluate my-run.sx --checkpoint s3://my-bucket/runs/my-run --follow --pass probe
It compiles the file under the checkpoint's selection, its variant and
every knob with the value experiment.json records, and the run's
--steps, --eval-every, --evaluate and --diagnostics overrides, so
the passes are the ones the run would have run and an unedited file shows
no change; --variant and --set are not taken.
It verifies the file against the checkpoint as a
resume does: a parameter that is new, gone, of
another shape or dtype, or trains where it was frozen is refused, naming
each one, and so is a :load file whose bytes changed (E-LOAD-001);
edited sources are allowed and printed under the checkpoint, one line
each. Then it builds the pass graphs alone, and the generator's when a
pass runs it; a with-loaded teacher is read from the files the
checkpoint pinned. The trained parameters come from params.safetensors;
the optimizer states and the training loader are never read. Each pass
reports into the run's own Metrics experiment, named by the slug
experiment.json records, at the checkpoint's step, with the series names
and metadata it has inline; a checkpoint without a slug is refused. The
experiment keeps describing the training run: its open carries the
checkpoint's manifest, every pass whatever --pass picked, and the
device kind and count the checkpoint's choices record. See metrics.
| flag | meaning |
|---|---|
--checkpoint <dir>|s3://bucket/prefix/step-<n> | evaluate that checkpoint |
--checkpoint <dir>|s3://bucket/prefix | a directory of step-<n> checkpoints: the newest complete one, or with --follow each as it appears |
--checkpoint latest | the same under --checkpoint-dir or SEXPGPU_CHECKPOINT_DIR |
--pass <name> | only this pass; repeatable. Every pass by default; an undeclared one is an error |
--follow | watch the directory: evaluate each new checkpoint once, in step order, until the run's last |
--poll <seconds> | how often --follow lists the directory; 60 |
--device, --metrics, --choices, --deterministic, --checkpoint-dir | as for run; SEXPGPU_DEVICES picks the GPU, its first ordinal, so an evaluation beside a training run can keep off the run's |
Without --follow, every selected pass runs on the one checkpoint,
whatever its :every. A machine with no GPU evaluates small passes with
--device cpu.
Choices. The checkpoint's recorded choices are taken when this is the
kind of device they were made on, and the evaluation is at
level 2: the CPU interpreter
repeats an inline CPU run's eval/* values bit for bit. On another kind
of device it chooses its own and is at level 3. --choices and
--deterministic replace either. The lowering report and the first
annotation say which.
Memory. On a GPU the memory plan holds the parameters, the loaded
teacher among them, the generator's phases and each pass's graph; no
optimizer state, gradient or training graph. A pass that does not fit
stops before it starts, E-MEM-001, as a run does; see
devices.
Following. --follow remembers each pass's evaluated steps in
evaluated.json beside the checkpoints (see
checkpoints), written after each
pass once its numbers have reached the metrics target, so a follower
restarted on any machine repeats no step it did and loses none it
marked. Numbers that do not arrive within 60 seconds stop the follower
with an error, the step undone. It
stops after the checkpoint whose experiment.json says "finished": true, which a run writes in the checkpoint of its last step. A follower started before
the first checkpoint waits for it.
What it never does. It writes no checkpoint, no status file, and never starts or finishes the experiment: the run's state in Metrics is the training run's. One evaluation can run beside the training run, or long after it.
Stopping. SIGTERM, SIGINT or SIGHUP stops a pass at its next
batch; that pass reports nothing, and in --follow its step stays undone
for a restart. A follower waiting for a checkpoint stops at once. The
process prints sexpgpu evaluate: interrupted by <signal> and exits
128 + n.
Output. On standard error: where the numbers go, the lowering report
on a GPU, evaluate: on <device>, determinism <level>, one evaluate: <checkpoint> line per checkpoint with any edited source under it, one
progress line per pass, and a done: block with each
pass's numbers.
| code | when |
|---|---|
0 | every pass asked for ran |
1 | refused or failed: diagnostics, an identity that does not match, no checkpoint, an undeclared --pass, a pass that did not fit |
2 | the command line was wrong: no file, no --checkpoint, a --poll that is not a positive number |
128 + n | stopped by signal n |
The pass named eval
The pass named eval is the one the summary, the status file, the progress
line and sweep's table report; a file without one reports its first pass.
Every observation from a pass carries pass metadata, so eval/loss from
eval and from probe are two series. See
metrics.
Errors
| code | when |
|---|---|
E-CONTRACT-013 | two passes share a name |
E-CONTRACT-014 | an evaluate or a curriculum in a file with no pass, so nothing would run it |
E-CONTRACT-003 | a pass's :prepare or :evaluate that is not a function, a defrun :evaluate other than :inline or :offline |
E-CONTRACT-018 | an offline run whose checkpoints do not land on a pass's :every |
Related: curriculum, the status file, checkpoints, determinism.