Benchmarks
At a glance
| Model | Recipe | Identity | Knowledge | Verdict |
|---|---|---|---|---|
| Spark | 1.5B MLX LoRA (rank 8) | 8/8 in weights | 8/8 factual | Identity genuinely learned, knowledge intact |
| Quark | 1.5B from scratch + SFT | 5/8 partially learned | 4/8, thin data | The real from-scratch effort: 35.8% HellaSwag, beats GPT-2 124M |
| Mini | 42M from scratch | — | Expect wrong answers | Pipeline proof, not a product |
Quark benchmarks
1.5B trained from scratch.
Spark benchmarks
Identity in the weights.
Mini benchmarks
42M from scratch.
Retired — Pulse & Pulse 2
| Model | Recipe | Identity | Knowledge | Verdict |
|---|---|---|---|---|
| Pulse | 1.5B fine-tune | 0/8 without system prompt | 8/8 factual | Identity came from the system prompt — superseded by Spark |
| Pulse 2 | 8B QLoRA (rank 16) | 0% without system prompt | 100% factual accuracy retained | Fine-tune was inert; branding came from the prompt |
Lattice