Tutorial
This tutorial rebuilds skycmd from an empty Cargo workspace to the games on the Play page. Every chapter follows the real code in the repository, with excerpts from the crates and the exact commands. Every number comes from a file in the repository (BENCH.md, runs/, data/STATS.md, docs/compat.md), and the chapters name it. A number nobody measured is marked "not measured".
What you end up with:
- Tiny. s1-pico has 202,305 parameters in a 545,280-byte file. s2-nano has 1,532,033 parameters in a 701,568-byte ternary file (
site/public/models/manifest.json). - Fast. One Flappy Drone decision, two typed questions plus the physics step, takes 694 µs for s1-pico and 501 µs for the s2-nano shape in WASM under Node 24 (
BENCH.md). - Typed. A model answers
choice,scoreandnoulquestions with a calibrated probability for every allowed option, in one forward pass, and a coded validator keeps the last word. - Plays games. After DAgger, s1-pico flies 200 of 200 pipes and rescues every survivor; s2-nano also matches the oracle's 0.87 Dodge completion. A fine-tuned s2-nano plays a Doom-like shooter at the oracle's level.
- Installs everywhere. The same
.gguffile runs in the browser, Node, Deno, Bun, Python, theskycmdCLI and server; the s1 files also run in llama.cpp and wllama (Install,docs/compat.md).
It also says what did not work: a model that crashed in every Flappy episode, an early-exit threshold that broke closed-loop play, a bigger model that did worse, and a smaller s2 that never learned to play.
The pipeline in one picture:
text
Hugging Face datasets ─┐ ┌─ sky-infer (native) ── sky-bench, sky-arena, skycmd
game oracles (sky-games)┼─ sky-data ─ sky-train ─ export ─ GGUF ─┤
GLM teacher (z.ai) ─────┘ (records) (mlx-rs) └─ sky-web (wasm) ── this siteFoundations
- Why tiny decision models: System One models, one forward pass, no generation.
- Setup on Apple Silicon: Rust, Metal for mlx-rs, wasm-pack, Node.
- Typed decisions: schema, unknown and the veto rule:
choice,score,noul, and why a coded validator has the last word. - Game cores and oracles: three deterministic games, their state text and their oracles.
Data
Model and training
- The encoder and the decision head
- Muon, WSD and the mlx-rs training loop
- Pretraining and decision training
- Calibration and evaluation
- The size ladder: s1-pico, s1-nano and s1-micro side by side, and why more parameters did not help.
Shipping
- GGUF export and int8 quantization
- A WASM inference engine with SIMD128
- Wiring the games
- Bonus: llama.cpp /v1/systemone and wllama: tested agreement within 4.7e-4 (llama-server) and 1.1e-4 (wllama).
- Scaling up and troubleshooting
Going further
- A Doom-like shooter: a fourth game, played from 14 measured numbers.
- The s2 architecture: two streams with a static cache, permutation equivariance,
[NUM]value features, early exits, ternary QAT, and why the exit threshold has to be calibrated in closed loop. - What DAgger taught us: from 0 to 200 Flappy pipes, round by round, for s1-pico, s1-nano and s2-nano.