# Checkpoints and resume

A run checkpoints to a directory or an `s3://` prefix, and continues from
the newest complete checkpoint with `--resume latest`. A signal stops it at
a step boundary with a checkpoint of that step, so a preempted run loses
nothing.

```bash
sexpgpu run my-run.sx --checkpoint-dir s3://my-bucket/runs/my-run --resume latest
```

## Checkpoints

- One directory per step, `step-<8 digits>`, under the checkpoint location
  (`--checkpoint-dir` or `SEXPGPU_CHECKPOINT_DIR`).
- Written every `defrun :checkpoint-every` steps and once more after the
  last step; without a location nothing is written.
- Four files ([below](#what-a-checkpoint-holds)); the last written is
  `experiment.json`, and a checkpoint counts only once it exists, so one cut
  off by a machine going away is passed over for the one before it.
- With `s3://bucket/prefix` the same files go to
  `s3://bucket/prefix/step-<n>/`: the two tensor files as concurrent
  multipart uploads, the others concurrently. Credentials are checked
  before the first step, so a run never finds out hours in that it cannot
  write; see [S3 credentials](https://sx.041.io/docs/environment.md#s3-credentials).
- Every checkpoint is an `annotation` in the metrics stream, with the
  `seconds` from the first tensor leaving the device to the last byte
  stored, and the status file's `checkpoint`.

## Resume

`--resume` or `SEXPGPU_RESUME` takes:

| value | resumes from |
|---|---|
| `<dir>` or `s3://bucket/prefix/step-<n>` | that checkpoint; one that is not there is an error |
| `latest` | the highest complete `step-<n>` under the checkpoint location, or a fresh start when there is none; needs a checkpoint location |

A location that cannot be listed is an error, never a fresh start. Put
`--resume latest` on the command line from the first launch; the same
command then starts, restarts and continues.

- **Refused:** a parameter that is new, gone, or another shape or dtype, or
  that trains where it was frozen or the reverse; a different device or node
  count (`E-DP-002`); a recorded choice this run cannot take
  (`E-CHOICE-001`); a `:load` file whose bytes are not the ones the
  checkpoint's run read (`E-LOAD-001`), told by the tensor and the file's
  size and version, its ETag on S3 or the SHA-256 of a local file:

  ```text
  E-LOAD-001: a loaded parameter's file is not the one the checkpoint's run read:
    model.teacher.q: q_proj.weight in s3://my-bucket/llama/model-00001-of-00002.safetensors, 4976698672 bytes, "9b2cf535f27731c974343645a3985328-149" -> ...
  restore that file, or start the run again
  ```
- **Pinned:** a resume takes the choices its checkpoint recorded, the
  stacking, the convolution engines and the recompute budget, instead of
  choosing again, so a run that moved to another machine keeps its
  summation order. `--choices fresh` lets it choose; see
  [determinism](https://sx.041.io/docs/determinism.md#pinning-a-run-to-recorded-choices).
- **Allowed, and recorded:** edited sources or other knobs. The
  `resumed from` annotation names each file and knob that moved, one line
  each. For example, a run resumed after a comment was added to its file,
  with `--set lr=0.25 --steps 6`:

  ```text
  resumed from /tmp/my-run/step-00000004
  source my-run.sx 91d77bdc76be -> b05c928e631c
  knob lr 0.5 -> 0.25
  knob steps 4 -> 6
  ```

  The same bytes at another path, a bundle's copy of a `:load` file, say
  `load <param path> <old> -> <new>`.

  A bundle's sources are its own paths and rewritten files, so resuming a
  bundled run's checkpoint from the unbundled file lists each of them.

A resumed run reports into the same metrics experiment; see
[events](https://sx.041.io/docs/events.md#a-resumed-run-is-the-same-experiment).

## What a checkpoint holds

A checkpoint promises two things: the run continues from it, and its
weights load into the same model somewhere else. It holds exactly that:

```text
params.safetensors   one tensor per parameter but a frozen one read from a
                     file, keyed by its dotted path
states.safetensors   one tensor per optimizer state, "<param path>/<state>"
state.json           step, stage, records, counters, the loader's position, seed
experiment.json      the source files and their sha256, the variant, every
                     knob with its value and source, each parameter's path,
                     shape, dtype, trainable and tags, and for a `:load`
                     parameter the file it was read from, pinned; each
                     parameter's state names, the counters, precision, seed,
                     device count, the sexpgpu version, and the choices the
                     run made, the run's slug, and whether it is the
                     checkpoint of the run's last step (`finished`)
```

The rule for [frozen parameters](https://sx.041.io/docs/models.md#frozen-and-pretrained-parameters):
one made by its initializer is in `params.safetensors` like any other, so
its random features load elsewhere and a resume takes them back. One read
from a file with `:load` is not copied: its `experiment.json` entry has a
`load` object naming the file (the shard, for an index), the tensor, and
the file's size and version, its ETag on S3 or `sha256:<hex>` of a local
file, and a resume reads that file again and refuses another one
(`E-LOAD-001`). A trainable parameter that started from a file has a `load`
too, and is in `params.safetensors` with what it has learned:

```json
{"path": "model.teacher.q", "shape": [4096, 4096], "dtype": "f32", "trainable": false, "tags": [],
 "load": {"uri": "s3://my-bucket/llama/model-00001-of-00002.safetensors", "tensor": "model.layers.0.self_attn.q_proj.weight",
          "size": 4976698672, "version": "\"9b2cf535f27731c974343645a3985328-149\""}}
```

Every [parameter root](https://sx.041.io/docs/models.md#roots) is in a checkpoint by the same
rule, under its own name: `model.*` first, then a second root such as a
`with-loaded` teacher as `teacher.*`, in `experiment.json`'s `params` and,
when it trains or was made by its initializer, in `params.safetensors`. A
frozen teacher read from a file is named, not copied: each of its
parameters is under `loaded` with the tensor its name function produced.

```console
loaded
  teacher.embed  teacher.embed.weight in teacher.safetensors, 2216 bytes, sha256:f2fb0b73...
  teacher.head   teacher.head.weight in teacher.safetensors, 2216 bytes, sha256:f2fb0b73...
params.safetensors  2 tensors
  model.embed  F32  [32, 8]
  model.head   F32  [8, 32]
```

`experiment.json` is a few kilobytes for any model and is written last.
The compiled graphs, the lowering and the kernels are not in a checkpoint: a resume compiles them again from the sources and the
binary. An exact replay of a compiled artifact is what a
[bundle](https://sx.041.io/docs/bundle.md) is for.

### What evaluate reads

[`sexpgpu evaluate`](https://sx.041.io/docs/evaluation.md#sexpgpu-evaluate) reads a checkpoint as
a resume does, less: `experiment.json`, to compile the file under its
selection, verify the parameters and the `:load` files, take its choices
and find the Metrics experiment by `slug`; `state.json`, for the step, the
stage and the counters a pass reads; and `params.safetensors`. It never
reads `states.safetensors`, never opens the training loader, and writes
nothing into a checkpoint. A frozen `:load` parameter is read from its
pinned file.

`evaluate --follow` keeps one file of its own beside the `step-<n>`
directories, `evaluated.json`: each pass's evaluated steps, rewritten
whole after each pass once that pass's numbers have reached the metrics
target (the collector's queue settled on them), to a bucket with the
checkpoint credentials. `latest` and `--resume` never read it; delete it
to evaluate again.

One case still repeats a pass: a follower that dies after its numbers
are delivered and before the file is rewritten evaluates that step
again when restarted, and Metrics keeps both points at that step, since
it deduplicates by message and the second run's messages are new. The
values are the same at level 2. The other order, marked and not
delivered, does not happen.

```json
{"eval":[250,500],"probe":[50,100,150,200,250,300,350,400,450,500]}
```

A follower stops after the checkpoint whose `experiment.json` has
`"finished": true`: a run writes it in the checkpoint of its last step
(`step` equal to `:steps`), the one it writes after its loop, or the last
`:checkpoint-every` one when that is the same step.

### Inspecting one

`sexpgpu checkpoint <dir|s3://bucket/prefix/step-<n>>` prints what a
checkpoint is without loading it: the identity from `experiment.json`, the
step from `state.json`, and every tensor's name, dtype and shape from the
two safetensors headers. An `s3://` location is read with the checkpoint
role's credentials, reading only the headers. A location
without `experiment.json` is not a complete checkpoint and is an error.

```console
$ sexpgpu checkpoint /tmp/my-run/step-00000004 | head -12
checkpoint /tmp/my-run/step-00000004
step       4
sexpgpu    0.1.0
variant    -
devices    1
precision  f32
seed       7
counters   tokens
sources
  <core>                    e4f1b2681fd7e21aa5cfe45f6f1b42618f70206f8af4becbf5adb1472edd09d9
  my-run.sx                 1a74f5546c852533510ab69ca07859527eb813001fc991db683c1ed1013f1117
  ...
```

Then the knobs, the choices, `loaded` with each `:load` parameter's file,
tensor, size and version, and each tensor file with one line per tensor; a
parameter's line ends with `frozen` when it does not train and with its
tags. A frozen parameter read from a file is under `loaded` and in no
tensor file.

```console
loaded
  model.teacher.q  model.layers.0.self_attn.q_proj.weight in s3://my-bucket/llama/model-00001-of-00002.safetensors, 4976698672 bytes, "9b2cf535f27731c974343645a3985328-149"
params.safetensors  33 tensors
  ...
  model.noise                          F32  [32]  frozen
```

### Loading one elsewhere

The safetensors files are the interchange format. A `bf16` parameter is
stored as `f32`, and `experiment.json`'s `dtype` says what it means:

```python
import json
from safetensors.numpy import load_file  # safetensors.torch.load_file for tensors

step = "/tmp/my-run/step-00000004"
identity = json.load(open(f"{step}/experiment.json"))
params = load_file(f"{step}/params.safetensors")   # {"model.net.embed.table": array, ...}
states = load_file(f"{step}/states.safetensors")   # {"model.net.embed.table/m": array, ...}
for declared in identity["params"]:
    if declared["path"] not in params:  # frozen and read from a file
        source = declared["load"]       # {"uri", "tensor", "size", "version"}
        continue
    assert list(params[declared["path"]].shape) == declared["shape"]
```

From a bucket, read the object's bytes and give them to
`safetensors.numpy.load`:

```python
import boto3
from safetensors.numpy import load

body = boto3.client("s3").get_object(
    Bucket="my-bucket", Key="runs/my-run/step-00000004/params.safetensors"
)["Body"].read()
params = load(body)
```

## Signals

`SIGTERM`, `SIGINT` (Ctrl-C) or `SIGHUP` (the terminal went away) stops a
run at the next step boundary:

1. the step in flight finishes, or an evaluation pass in flight stops at
   its next batch, reports nothing and is run again by the resume; on
   several GPUs the pass finishes, since the ranks look for a signal only
   together. No further step is taken. A step whose
   [generator](https://sx.041.io/docs/generate.md) is still making its records stops before its
   next trip and is abandoned: no parameter has changed, its batches are
   given back and its counters undone, so it is the step before it that
   is kept, and the resume takes the abandoned step again, to the bits of
   a run never stopped. On several GPUs it finishes;
2. the metrics get `interrupted by <signal> at step <n>`, with
   ` while generating` after an abandoned step or pass;
3. a checkpoint of that step is written when there is a location;
4. the stream is drained without `done`, and the status file reads
   `interrupted`;
5. the process prints `sexpgpu run: interrupted by <signal> at step <n>`
   and exits `128 + n`: `129`, `130`, `143`.

A second `SIGTERM` or `SIGINT` ends the process at once, without any of it.
A `SIGHUP` never does, because a machine shutting down sends `SIGTERM` and
`SIGHUP` together. A run killed outright (`SIGKILL`, a second signal, the
machine gone) writes nothing more: its status file keeps reading `running`,
and its resume repeats the steps after its last checkpoint.

A spot VM's shutdown or a job runner's cancel gives the process a `SIGTERM`
and some seconds; a checkpoint larger than that allows is passed over. See
[preemptible jobs](https://sx.041.io/docs/bundle.md#preemptible-jobs).

## The finite guard

`(defrun ... :guard-finite true)` checks every proposed parameter and
optimizer state for NaN and infinity on the device before any of the step
is committed: one reduction and one downloaded number per tensor. When one
is not finite the run stops with the step, the parameter and the state in
the error, which the status file and a last `failed: <error>` annotation
repeat. Nothing of that step is committed, so the last checkpoint is intact
and a resume from it replays the step.

```console
sexpgpu run: step 2: the proposed parameter of model.b has 1 values that are not finite; nothing was committed
```

It is off by default because it adds a reduction and a download per
parameter per step.

Related: [run](https://sx.041.io/docs/run.md), [bundle](https://sx.041.io/docs/bundle.md), [the status file](https://sx.041.io/docs/status.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
