# Models and parameters

A model is a set of parameters plus the function that uses them.
`defmodel` defines a constructor; calling it creates fresh parameters and
returns the model; applying a model is calling it.

```lisp
(defmodel linear (in out &key (bias true) (init-std nil))
  (defparam weight (normal [in out] :std (or init-std (/ 1.0 (sqrt in)))))
  (when bias (defparam b (zeros [out])))
  (lambda (x) (if bias (+ (matmul x weight) b) (matmul x weight))))

(defvar model (linear 32 8))
```

## defmodel

- The body is ordinary Lisp: `let` sub-models, loop, branch on shapes.
- The body must end in a function, a builtin or another model
  (`E-PARAM-004`); return `(lambda (x) ...)` last.
- Each call to a constructor creates its own parameters. Importing a
  definition imports no state.
- The contract's `model` must be a model value, the result of calling a
  constructor; see [the run file](https://sx.041.io/docs/run-file.md). Any other model bound at
  the top of the run file has parameters too; see [roots](#roots).
- Weights are stored `[in, out]`: `(matmul x weight)`. A PyTorch
  `Linear` stores the transpose.

## defparam

`(defparam name initializer &key tags numerics trainable load)` declares
persistent state, inside a model only (`E-PARAM-002`).

- The initializer is a closed graph of its own: shapes, constants and
  random nodes, never another value (`E-PARAM-003`). The parameter's shape
  and dtype are its output's, so `(normal [in out] :std s)` says everything
  once.
- Random initializers fold in `defrun :seed`; gensym names never affect
  the initial values.

| keyword | value | effect |
|---|---|---|
| `:tags` | a list or vector of names | `(select :tag :name)` finds the parameter; see [optimizer groups](https://sx.041.io/docs/optimizer-groups.md) |
| `:numerics` | `:high` | under `:bf16` the parameter stays an `f32` master weight read without a cast down; see [numerics](https://sx.041.io/docs/numerics.md) |
| `:trainable` | a boolean, `true` by default | `false` freezes it; see [below](#frozen-and-pretrained-parameters) |
| `:load` | `[file tensor]` | starts it from `tensor` of a safetensors `file` instead of the initializer; see [below](#frozen-and-pretrained-parameters) |

All evaluate their argument expressions. Any other option, or a
`:numerics` other than `:high`, is `E-PARAM-006`; a `:trainable` that is not
a boolean, or a `:load` that is not a file and a tensor name, is
`E-PARAM-007`. [`with-loaded`](#with-loaded) gives `:trainable` and `:load`
to every parameter a model declares.

## Frozen and pretrained parameters

```lisp
(defmodel frozen-attention (width weights prefix)
  (defparam q (zeros [width width])
    :trainable false
    :load [weights (str prefix "q_proj.weight")])
  (lambda (x) (matmul x (transpose q [1 0]))))

(defvar teacher
  (frozen-attention 4096 "s3://my-bucket/llama/model.safetensors.index.json"
                    "model.layers.0.self_attn."))
```

- `:trainable false` freezes a parameter: no gradient reaches it and it
  joins no optimizer group, so a selector never matches it.
- `:load [file tensor]` starts a parameter from one tensor of a safetensors
  file. The file is an `s3://` URI, read with the data credentials, or a
  local path, relative to the experiment file as a loader path is. A
  `.index.json` names each tensor's shard, so a sharded Hugging Face
  checkpoint is loaded by its index and its own tensor names.
- The initializer still states the shape and dtype. The tensor must have
  that shape; `BF16` and `F32` tensors load into `:f32` or `:bf16`
  parameters. Any other, or a file whose header does not describe its
  bytes, stops the run before its first step.
- A Hugging Face checkpoint keeps PyTorch's `[out in]` layout; multiply by
  its transpose, as above.
- The two are independent. `:load` alone fine-tunes from the file;
  `:trainable false` alone is a fixed random feature.
- A frozen parameter is stored in the dtype its readers read: `bf16` when
  every use of an `f32` one is cast down, under `:bf16` or inside a
  `bf16` region. See [numerics](https://sx.041.io/docs/numerics.md#parameters).
- A checkpoint holds a frozen parameter made by its initializer, like any
  other. A frozen parameter read from a file is not copied: the
  checkpoint's `experiment.json` names the file, the tensor, and the file's
  size and version, and a resume reads it again. A resume refuses a file
  whose bytes changed (`E-LOAD-001`), and a parameter that trains where it
  was frozen or the reverse. See [checkpoints](https://sx.041.io/docs/checkpoints.md#what-a-checkpoint-holds).
- A [bundle](https://sx.041.io/docs/bundle.md) copies a local `:load` file and every shard its
  index names; an `s3://` file is read where the bundle runs.

### with-loaded

A pretrained model has hundreds of parameters, and its constructor should
not need a `:load` on each. `with-loaded` gives them one:

```lisp
(defun olmo3-tensor-name (path)
  ;; "teacher.blocks.3.attn.q" -> "model.layers.3.self_attn.q_proj.weight"
  ...)

(defvar teacher
  (with-loaded ("s3://my-bucket/olmo3/model.safetensors.index.json" olmo3-tensor-name)
    (olmo3 :layers 32 :width 4096)))
```

`(with-loaded (file name &key (trainable false)) body...)` evaluates its
body, and every `defparam` that executes meanwhile, in any constructor the
body calls, loads from `file`, a `.safetensors` or an `.index.json`, local or
`s3://`, as `:load` takes it.

- `name` is a function of one argument. It is called with each parameter's
  [path](#paths), the final one that checkpoints and `explain` show, root
  included (`teacher.blocks.3.attn.q`), once the file is evaluated, and
  returns the tensor's name in the file. It runs after the whole file is
  evaluated, so a top-level name it reads has the value it has then, not
  the one it had when the `with-loaded` ran. Anything but a string is
  `E-PARAM-007`, naming the path and the value; an error inside it says
  which path it was called for. A tensor missing from the
  file, of another shape, or of a dtype that does not load stops the run
  before its first step with the path and the tensor's name.
- The parameters are frozen: `:trainable false` is the default inside.
  `(with-loaded (file name :trainable true) ...)` fine-tunes them.
- A parameter's own `:load` or `:trainable` wins, each on its own: a
  `defparam` with `:load` reads its own file and still freezes, one with
  `:trainable true` still loads from `file`.
- Forms nest, and the innermost wins: a `with-loaded` inside a constructor
  called under another one loads its parameters from its own file.
- Only a `defparam` that executes while the body is evaluated is affected:
  a constructor called later, from a function the body returned, is not.
- A header that is not a file, a function and an optional `:trainable`
  boolean is `E-PARAM-007`.
- A [bundle](https://sx.041.io/docs/bundle.md) copies a local file and rewrites its path when the
  header writes it as a string, as it does a `:load`'s, and `sexpgpu new`
  repoints it. A path the header computes is not rewritten.

## Roots

Every model bound to a name at the top of the run file, defined there with
`defvar` or imported by `require`, is a parameter root under that name, as
`model` is. Its parameters exist, have paths under its name, are read by any
graph, train unless frozen and are checkpointed:

```lisp
(defvar teacher (with-loaded ("teacher.safetensors" hf-name) (bigram 32 8)))
(defvar model (bigram 32 8))

(defun prepare (batch)
  (let ((ids (field batch :ids)))
    (list :inputs ids :targets ids :teacher (teacher ids))))
```

The parameters are `model.*` first, then each other root in the order the
file binds it, imports first.

- A model has one root. The same model bound to two top-level names, or a
  model reachable from two roots, is `E-PARAM-008`, naming both.
- Only a model bound by itself is a root. A list or vector of models bound
  at the top level is not one, and its models are not either; bind each to
  its own name, or build them inside a `defmodel`.
- A model built inside a `let`, a function, a top-level list or a library
  module and not bound by itself at the top of the run file has no root. A
  graph that reads one of its parameters is `E-PARAM-009`: bind it with
  `defvar` at the top level, or build it inside the `defmodel` that uses it.

## Paths

A parameter's canonical path comes from binding names:

- the root is the top-level name the model is bound to, normally `model`,
  whatever `named` or an earlier `let` called it;
- inside a `defmodel`, a model bound by `let` takes its binding name;
- a list or vector of models takes one index per element;
- a `defparam` takes its own name.

```lisp
(let ((blocks (repeat 12 (lambda (i) (block dim heads))))) ...)
;; model.blocks.3.attn.q.weight
```

- `(named "segment" value)` overrides the segment of a model or parameter
  (`E-PARAM-005` on anything else).
- A model never bound to a name gets `anon.N` and the warning
  `W-PARAM-001`.
- Gensym bindings never name models; a parameter declared with a gensym
  gets a stable `param.N` within its model.

Paths are what [selectors](https://sx.041.io/docs/optimizer-groups.md#selectors) match, what
[`explain`](https://sx.041.io/docs/explain.md) lists and what checkpoints store. Read them with
`sexpgpu explain` before writing a selector.

## Weight tying

Tying is value sharing: use the same parameter, or the same model, in two
places and it is one parameter, updated once.

```lisp
(defmodel tied-lm (vocab dim)
  (defparam table (normal [vocab dim] :std 0.02))
  (lambda (ids) (matmul (index table ids) (transpose table [1 0]))))
```

`model.table` is one parameter; its gradient sums both uses.

Related: [sexpgpu/nn](https://sx.041.io/docs/nn.md), [numerics](https://sx.041.io/docs/numerics.md), [tensor operations](https://sx.041.io/docs/tensors.md).

---

SexpGPU documentation. Every page: https://sx.041.io/llms.txt
