S-exp GPU
All pages
Docs · Models and optimizersMarkdown

Models and parameters

A model is a set of parameters plus the function that uses them. defmodel defines a constructor; calling it creates fresh parameters and returns the model; applying a model is calling it.

(defmodel linear (in out &key (bias true) (init-std nil))
  (defparam weight (normal [in out] :std (or init-std (/ 1.0 (sqrt in)))))
  (when bias (defparam b (zeros [out])))
  (lambda (x) (if bias (+ (matmul x weight) b) (matmul x weight))))

(defvar model (linear 32 8))

defmodel

  • The body is ordinary Lisp: let sub-models, loop, branch on shapes.
  • The body must end in a function, a builtin or another model (E-PARAM-004); return (lambda (x) ...) last.
  • Each call to a constructor creates its own parameters. Importing a definition imports no state.
  • The contract's model must be a model value, the result of calling a constructor; see the run file. Any other model bound at the top of the run file has parameters too; see roots.
  • Weights are stored [in, out]: (matmul x weight). A PyTorch Linear stores the transpose.

defparam

(defparam name initializer &key tags numerics trainable load) declares persistent state, inside a model only (E-PARAM-002).

  • The initializer is a closed graph of its own: shapes, constants and random nodes, never another value (E-PARAM-003). The parameter's shape and dtype are its output's, so (normal [in out] :std s) says everything once.
  • Random initializers fold in defrun :seed; gensym names never affect the initial values.
keywordvalueeffect
:tagsa list or vector of names(select :tag :name) finds the parameter; see optimizer groups
:numerics:highunder :bf16 the parameter stays an f32 master weight read without a cast down; see numerics
:trainablea boolean, true by defaultfalse freezes it; see below
:load[file tensor]starts it from tensor of a safetensors file instead of the initializer; see below

All evaluate their argument expressions. Any other option, or a :numerics other than :high, is E-PARAM-006; a :trainable that is not a boolean, or a :load that is not a file and a tensor name, is E-PARAM-007. with-loaded gives :trainable and :load to every parameter a model declares.

Frozen and pretrained parameters

(defmodel frozen-attention (width weights prefix)
  (defparam q (zeros [width width])
    :trainable false
    :load [weights (str prefix "q_proj.weight")])
  (lambda (x) (matmul x (transpose q [1 0]))))

(defvar teacher
  (frozen-attention 4096 "s3://my-bucket/llama/model.safetensors.index.json"
                    "model.layers.0.self_attn."))
  • :trainable false freezes a parameter: no gradient reaches it and it joins no optimizer group, so a selector never matches it.
  • :load [file tensor] starts a parameter from one tensor of a safetensors file. The file is an s3:// URI, read with the data credentials, or a local path, relative to the experiment file as a loader path is. A .index.json names each tensor's shard, so a sharded Hugging Face checkpoint is loaded by its index and its own tensor names.
  • The initializer still states the shape and dtype. The tensor must have that shape; BF16 and F32 tensors load into :f32 or :bf16 parameters. Any other, or a file whose header does not describe its bytes, stops the run before its first step.
  • A Hugging Face checkpoint keeps PyTorch's [out in] layout; multiply by its transpose, as above.
  • The two are independent. :load alone fine-tunes from the file; :trainable false alone is a fixed random feature.
  • A frozen parameter is stored in the dtype its readers read: bf16 when every use of an f32 one is cast down, under :bf16 or inside a bf16 region. See numerics.
  • A checkpoint holds a frozen parameter made by its initializer, like any other. A frozen parameter read from a file is not copied: the checkpoint's experiment.json names the file, the tensor, and the file's size and version, and a resume reads it again. A resume refuses a file whose bytes changed (E-LOAD-001), and a parameter that trains where it was frozen or the reverse. See checkpoints.
  • A bundle copies a local :load file and every shard its index names; an s3:// file is read where the bundle runs.

with-loaded

A pretrained model has hundreds of parameters, and its constructor should not need a :load on each. with-loaded gives them one:

(defun olmo3-tensor-name (path)
  ;; "teacher.blocks.3.attn.q" -> "model.layers.3.self_attn.q_proj.weight"
  ...)

(defvar teacher
  (with-loaded ("s3://my-bucket/olmo3/model.safetensors.index.json" olmo3-tensor-name)
    (olmo3 :layers 32 :width 4096)))

(with-loaded (file name &key (trainable false)) body...) evaluates its body, and every defparam that executes meanwhile, in any constructor the body calls, loads from file, a .safetensors or an .index.json, local or s3://, as :load takes it.

  • name is a function of one argument. It is called with each parameter's path, the final one that checkpoints and explain show, root included (teacher.blocks.3.attn.q), once the file is evaluated, and returns the tensor's name in the file. It runs after the whole file is evaluated, so a top-level name it reads has the value it has then, not the one it had when the with-loaded ran. Anything but a string is E-PARAM-007, naming the path and the value; an error inside it says which path it was called for. A tensor missing from the file, of another shape, or of a dtype that does not load stops the run before its first step with the path and the tensor's name.
  • The parameters are frozen: :trainable false is the default inside. (with-loaded (file name :trainable true) ...) fine-tunes them.
  • A parameter's own :load or :trainable wins, each on its own: a defparam with :load reads its own file and still freezes, one with :trainable true still loads from file.
  • Forms nest, and the innermost wins: a with-loaded inside a constructor called under another one loads its parameters from its own file.
  • Only a defparam that executes while the body is evaluated is affected: a constructor called later, from a function the body returned, is not.
  • A header that is not a file, a function and an optional :trainable boolean is E-PARAM-007.
  • A bundle copies a local file and rewrites its path when the header writes it as a string, as it does a :load's, and sexpgpu new repoints it. A path the header computes is not rewritten.

Roots

Every model bound to a name at the top of the run file, defined there with defvar or imported by require, is a parameter root under that name, as model is. Its parameters exist, have paths under its name, are read by any graph, train unless frozen and are checkpointed:

(defvar teacher (with-loaded ("teacher.safetensors" hf-name) (bigram 32 8)))
(defvar model (bigram 32 8))

(defun prepare (batch)
  (let ((ids (field batch :ids)))
    (list :inputs ids :targets ids :teacher (teacher ids))))

The parameters are model.* first, then each other root in the order the file binds it, imports first.

  • A model has one root. The same model bound to two top-level names, or a model reachable from two roots, is E-PARAM-008, naming both.
  • Only a model bound by itself is a root. A list or vector of models bound at the top level is not one, and its models are not either; bind each to its own name, or build them inside a defmodel.
  • A model built inside a let, a function, a top-level list or a library module and not bound by itself at the top of the run file has no root. A graph that reads one of its parameters is E-PARAM-009: bind it with defvar at the top level, or build it inside the defmodel that uses it.

Paths

A parameter's canonical path comes from binding names:

  • the root is the top-level name the model is bound to, normally model, whatever named or an earlier let called it;
  • inside a defmodel, a model bound by let takes its binding name;
  • a list or vector of models takes one index per element;
  • a defparam takes its own name.
(let ((blocks (repeat 12 (lambda (i) (block dim heads))))) ...)
;; model.blocks.3.attn.q.weight
  • (named "segment" value) overrides the segment of a model or parameter (E-PARAM-005 on anything else).
  • A model never bound to a name gets anon.N and the warning W-PARAM-001.
  • Gensym bindings never name models; a parameter declared with a gensym gets a stable param.N within its model.

Paths are what selectors match, what explain lists and what checkpoints store. Read them with sexpgpu explain before writing a selector.

Weight tying

Tying is value sharing: use the same parameter, or the same model, in two places and it is one parameter, updated once.

(defmodel tied-lm (vocab dim)
  (defparam table (normal [vocab dim] :std 0.02))
  (lambda (ids) (matmul (index table ids) (transpose table [1 0]))))

model.table is one parameter; its gradient sums both uses.

Related: sexpgpu/nn, numerics, tensor operations.