# Metrics

A run reports in two independent channels. Metrics are the time series:
one event per number, sent as it happens to a file, a pipe or
[Metrics by 041](https://metrics.041.io/). The [status file](https://sx.041.io/docs/status.md) is
the single current state, rewritten in place. Consume the metrics stream;
poll the status file.

## Targets

`SEXPGPU_METRICS`, or `--metrics`, read once at run start:

| target | where numbers go |
|---|---|
| unset | nowhere |
| `jsonl:<path>` | one JSON object per line, appended; parent directories are created |
| `stdout` | the same events on standard output, unbuffered, one per line |
| `o41` | [Metrics by 041](https://metrics.041.io/) |
| `o41:<folder-id>` | the same, in that folder |

`o41` needs `O41_METRICS_API_KEY`; the folder is
`O41_METRICS_FOLDER_ID`, and one named in the target wins. Anything else
is an error at start. The run prints where its numbers go before it trains
(`metrics: /tmp/my-run.jsonl`, or the experiment's URL), and the same string
is in the summary, the status file and `sweep`'s table. With `o41` the
run waits up to five seconds for the experiment id, then trains without a
URL to print.

Events are delivered in order, from one thread, so a run that reports
faster than Metrics by 041 accepts falls behind. When a run ends or
stops it waits up to thirty seconds for the queue, so a finished run
catches up, but a preempted one whose machine goes first loses what was
still queued, its `interrupted` and `checkpoint` annotations included; the
checkpoint itself is written before the wait.

`stdout` is the plugin interface: `sexpgpu run ... | my-forwarder` sends
events anywhere without the backend knowing. Standard output then carries
nothing but events, unless `--json` is also given: then its last line is the
[summary](https://sx.041.io/docs/status.md#the-summary), of kind `summary`, which is not an event
and has no `ts_ms`. A forwarder filters on `kind`. The event format is in
[JSONL events](https://sx.041.io/docs/events.md).

## Reporting from the experiment

`(metric "name" scalar :key value ...)` records one scalar under a name,
takes the keywords after it as metadata, and returns the value, so it wraps
an expression. It works anywhere a step runs:

| where | what it becomes |
|---|---|
| `objective`, `prepare` | averaged over a step's microbatches, emitted at that step, beside `train/objective` |
| `evaluate` | averaged over the batches of a pass, emitted at the step it ran, tagged with the pass, offered to the curriculum by name |
| the model body | observed in every graph the model is applied in: a training-step average and one per evaluation pass, under one name told apart by `split` and `pass`; a library model reports the same way |
| an optimizer body | refused, `E-CONTRACT-011`: use a [diagnostic](https://sx.041.io/docs/writing-diagnostics.md#optimizer-diagnostics) |
| `curriculum` | not read; the curriculum is a decision, not a report |

```lisp
(metric "probe/logit-rms" (sqrt (mean (* logits logits))) :layer "out")
```

A number about how the model is doing rather than how well, that nobody
needs on every microbatch of every run, is a [diagnostic](https://sx.041.io/docs/diagnostics.md):
it costs nothing until selected.

Metadata values are strings, keywords, numbers or booleans, recorded as
strings; a tensor is refused, because metadata is fixed at compile time.
Metadata is capped at 32 keys, 128-byte keys, 512-byte values and 4 KiB
of canonical JSON. The compiler sizes every `metric` exactly,
with the runtime's keys and the longest stage name counted, and refuses
one over the limit as `E-CONTRACT-012` rather than letting the run stop at
its first event.

## Offline evaluation

A pass [`sexpgpu evaluate`](https://sx.041.io/docs/evaluation.md#sexpgpu-evaluate) runs from a
checkpoint reports into the training run's experiment, found by the slug
the checkpoint records, at the checkpoint's step, under the same name and
metadata as inline: `eval/loss` with `pass` and `split`, then
`curriculum/stage` and `runtime/eval_seconds`. The series line up with
training as if the pass had run in the loop. Before its first numbers,
each `evaluate` process adds one annotation at that step naming the device
it evaluates on and its [level](https://sx.041.io/docs/determinism.md#offline-evaluation):

```json
{"step":300,"kind":"annotation","annotation":"evaluated offline on cpu: determinism 3: chosen here, ...","metadata":{"device":"cpu","level":"3: chosen here, ...","checkpoint":"s3://my-bucket/runs/my-run/step-00000300"}}
```

Its `metadata` also carries `changed`, the source and knob lines a resume
would print, so an evaluation of an edited file is visible in Metrics.

It opens the experiment with the `open` event, built from the checkpoint
so that it describes the training run as the run's own did: the
checkpoint's manifest, every pass, the run configuration, and the device
kind and count its recorded choices name (`patterns` and the node count,
which a checkpoint does not record, are left out). It never sends `start`
or `done`, so the experiment's state is the training run's. A signal stops a pass with nothing of it reported. An offline run
itself reports no pass at all.

## Runtime metadata

The runtime adds its own keys to every event of a series:

| key | on | value |
|---|---|---|
| `split` | every metric | `train` from the training graph and counters, `eval` from every pass and `curriculum/stage` |
| `pass` | every metric of an evaluation pass, and its `curriculum/stage` | the `defeval` name, `eval` or `probe` |
| `stage` | every metric, when the training loader has several stages | the stage it was measured in |
| `param` | every optimizer diagnostic | the update group its values folded over |

A series is its name and its metadata together, so `train/loss` from the
`easy` and the `hard` stage are two series. An experiment that writes
`:split`, `:pass` or `:stage` itself keeps its own value.

## Experiment limits

Metrics by 041, used by the `o41` targets, has these limits:

| what | maximum |
|---|---|
| distinct series names, meaning metric names, per experiment | 4096 |
| unique metadata combinations per series name in one experiment | 4096 |

The metadata limit counts distinct complete metadata maps, including
runtime-added keys such as `split`, `pass`, `stage`, and `param`. Repeated
points with the same name and metadata reuse the same combination. These
limits apply across the whole experiment, including resumed runs and offline
evaluation.

## Metric names

| name | what it is |
|---|---|
| whatever `metric` was called with | an observation, averaged over a step's microbatches or a pass's batches |
| `diagnostics/<name>` | a selected [diagnostic](https://sx.041.io/docs/diagnostics.md), on sampled steps, folded by its reducer |
| `train/objective` | the objective's own value, whatever its reduction and name |
| `counter/<name>` | a declared [counter](https://sx.041.io/docs/data.md#counters)'s running total, every step |
| `rate/<counter>` | what the counter grew by this step, per second of `runtime/step_seconds` |
| `runtime/step_seconds` | the whole step, first batch to end of update, device synchronized, evaluation excluded |
| `runtime/train_seconds` | the sum of `runtime/step_seconds` so far in this process |
| `runtime/eval_seconds` | one evaluation pass, synchronized, with its `pass` |
| `runtime/peak_bytes` | the device allocator's high-water mark of live bytes, plus the private execution cache's peak on CUDA; absent on the CPU |
| `curriculum/stage` | the stage index an evaluation ran in, with every evaluation |

Every run reports the runtime series with nothing to select. The runtime
knows no domain, so `(counter :tokens ...)` gives `rate/tokens`, tokens
per second. On the CPU timings are wall time; under data parallel they are
rank zero's. `train/loss` and `eval/loss` mean something to the runtime; see
[objective](https://sx.041.io/docs/objective.md#two-names-the-runtime-reads).

Related: [JSONL events](https://sx.041.io/docs/events.md), [diagnostics](https://sx.041.io/docs/diagnostics.md),
[the status file](https://sx.041.io/docs/status.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
