# doctor

`sexpgpu doctor` checks whether a run, or an
[evaluation](https://sx.041.io/docs/evaluation.md#sexpgpu-evaluate) from checkpoints, can start on this machine: one line
per check on standard output, exit `0` or `1`. Run it first on a new
machine.

```console
$ sexpgpu doctor
sexpgpu      0.1.0 /usr/local/bin/sexpgpu
cuda build   this binary has no --features cuda; SEXPGPU_DEVICE=cuda would refuse
metal        SEXPGPU_DEVICE=metal needs macOS and a build with --features metal
gpu          none (nvidia-smi is not on PATH or found no device)
patterns     none without --features cuda
env          no SEXPGPU_* variable is set; the defaults are cpu and no metrics
metrics key  O41_METRICS_API_KEY not set; SEXPGPU_METRICS=o41 would refuse
s3           no AWS_* or SEXPGPU_*_AWS_* variable is set; an s3:// loader or checkpoint would refuse
disk         72Gi free in the working directory
choices      measured and remembered: 0 choices in /home/me/.cache/sexpgpu/choices.json, the cost model's constants for 0 devices; a resume takes its checkpoint's
run          ok: a run or an evaluation can start here
```

On a machine with a GPU a `floor` line follows `gpu`, saying whether the
driver and the compute capability meet the
[CUDA requirements](https://sx.041.io/docs/devices.md#the-cuda-device), and the CUDA binary adds a
`devices` line with the CUDA ordinals a run would use and a `patterns` line
saying whether the pattern kernels are on. A `nodes`
line shows where a run over [several nodes](https://sx.041.io/docs/devices.md#several-nodes) puts
this process. With `SEXPGPU_CHECKPOINT_DIR` set, a `checkpoints` line
names the newest complete checkpoint there, what `evaluate --checkpoint
latest` and `run --resume latest` would read. The `floor` line is informational: a card or driver below
the floor does not make `doctor` exit `1`, but a run refuses the card when
it opens the device, so read it before trusting `run ok`. Keys are
reported as set or not, never by value.

A Metal-enabled macOS binary adds a `metal` line with the Apple GPU name
or the reason it cannot open the GPU.

## When it exits 1

When any of these holds, each of which would stop a run before its first
step:

- an unreadable `SEXPGPU_*` configuration;
- `SEXPGPU_DEVICE=cuda` in a binary without CUDA, or with no GPU;
- `SEXPGPU_DEVICE=metal` without a Metal-enabled macOS binary or an Apple GPU;
- an `SEXPGPU_DEVICES` that does not parse or names an ordinal that is not
  visible, `E-DP-001`;
- a node topology that does not parse, `E-DP-005`;
- an `s3://` checkpoint location without a full set of credentials;
- a local `SEXPGPU_RESUME` that is not a directory;
- an `SEXPGPU_CHECKPOINT_DIR` that cannot be listed, which `evaluate` and
  `run --resume latest` would read;
- an `SEXPGPU_CHOICES` that names no recorded choices, `E-CHOICE-002`.

The `choices` line says how a run would make its
[choices](https://sx.041.io/docs/determinism.md): pinned, none measured, or measured and
remembered in the cache it counts.

Device availability refusals apply to the selected `SEXPGPU_DEVICE`. `doctor` reads
the same [environment](https://sx.041.io/docs/environment.md) a run would.

Related: [devices](https://sx.041.io/docs/devices.md), [install](https://sx.041.io/docs/install.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
