Research Notes · The series so far
Three weeks, measured
Every result we have published, on one page — a dated timeline of the cadence, and a map of the whole GoM stack. Nothing here is new; every number links back to the note that reports it.
W4M Research · July 2026 · the series so far
14 notes
published research notes between Jun 30 and Jul 18, 2026 — every figure below links to the note that reports it
6 in a day
the founding drop: Notes 1–6 all went out on Jun 30, 2026 — then the cadence kept going
+57.5 points
just published in Note 14 — the frozen Nemotron Nano's GSM8K number climbed again under a leaner brain: 25.0 → 82.5% (n=200). The pace is accelerating.
How to read this page
Two pictures. First, a timeline: what landed, when, in the order it published — the point is the density of the cadence, not any single line. Second, a map of the stack: the small models we train ourselves, the brains that bolt onto someone else's frozen model, and the innovations around both. Bolded numbers are all previously published; each card links to its source note.
Visual 1 — Three weeks, measured
A vertical timeline of the series. Same-day clusters are drawn together on purpose: on Jun 30 six notes went out at once, and the drumbeat has not stopped since.
Jul 8, 20263 notes
Note 9 · Reasoning at depth
The full 30-disk tower — 1,073,741,823 moves, every one exact, in ~9 minutes on a laptop CPU.
Note 10 · Exploratory
0 claims of machine consciousness — but a self-model that predicts its own success 96–98% of the time.
Note 11 · Precision
Sudoku's hardest tier, solved: 4,865 / 4,865 (100.00%) from one 400 KB model, beating Kona's 96.2%.
Honest note
This page adds no new claims. Every bolded figure is lifted from a note already published above, and each card carries the link so you can check it in context — we do not put figures on this site before the note that reports them.
Visual 2 — The GoM stack
The same body of work, arranged by what it is rather than when it shipped: the models we train ourselves (small to large), the brains that attach to a frozen third-party model, and the innovations that surround both.
Column 1
GoM core models
Models we train from scratch — small to large.
98K-param Sudoku specialist · 400 KB
Solves the complete 17-clue tier — 4,865 / 4,865 (100.00%), independently audited — from one 400 KB checkpoint, past Kona's 96.2%.
Note 11 →
~37M reasoning core · Tower-of-Hanoi class
Trained only on 3–8 disks, it produced the full 30-disk tower — 1,073,741,823 moves, every one exact, in ~9 min on a laptop CPU.
Note 9 →
ARC program-writer cores
Write candidate programs, verify them against a puzzle's own examples, and repair — on a held-out fresh-transfer probe, 96.7% of tasks yielded machine-verified programs.
Note 13 →
Scaling family · 30M→962M
Five sizes, one recipe, on a clean power law — exponent β = 0.245 (R² = 0.938), zero training divergences. Larger bases in training.
Note 1 →
Column 2
GoM-Brain variants
A small brain that attaches to a frozen LLM and lifts it — no base weights touched.
376M brain → Nemotron Nano (3B-active, hybrid Mamba-MoE)
25.0 → 37.5 → 69.5 → 82.5%
GSM8K on a completely frozen model, measured three times in one week — each better than the last: +12.5 (first onboarding) → +44.5 (a deeper checkpoint-selection instrument) → +57.5 points (a slimmer brain variant, n=200). Same 376M-class brain, same frozen base — the method got sharper, not the model bigger.
Note 14 →
Brain → 120B Nemotron Super
53.5 → 76.5%
GSM8K, +23 points, n=200 both arms, no weights touched — interim, formal verdict pending.
Note 13 →
27B-host brain · on-device
+26.5 points
Our strongest deployed result to date — a 27B base lifted on a laptop at 4-bit, no cloud in the loop.
Note 13 →
Do-no-harm confidence gate
The brain engages where it helps and is measured harm-free across a standard benchmark row (ARC-E / ARC-C / PIQA / HellaSwag) — a capability, not a pledge.
Note 13 →
Automated onboarding
Point it at a frozen model and it attaches a brain unattended — the Nano brain above trained overnight on a single GPU. Any base, one afternoon.
Note 13 →
Code-focused variants In testing
Code-specialized brains are in the lab. Numbers when they clear the bar — not before.
Column 3
Beyond the models
The innovations around the models — memory, sharing, deployment.
On-device assistant (GoM-Siri)
42 → 68%
Our on-device assistant work: 42 → 68% on a hard 714-question test, 0 confidently-wrong answers, every query kept on the phone.
Note 6 →
Sleep, consolidation & memory rekeying
It sleeps and it remembers — a spaced schedule holds 81% of lessons at day ten versus 32% without it.
Note 7 →
Portable episodic memory
Memories live inside the model file, not a server — so experience transfers as a file: a fresh unit goes 0 → 100% on a donor's material with no retraining.
Note 8 →
Swarm experience-sharing
45.6 → 100%
Three brains, three lives, one overnight sync — every member woke up perfect on the full exam, replicated ×3.
Note 8 →
Self-improvement loop
It taught itself from 87 → 100% with zero human data, behind a verifier that gates its own generated examples.
Note 7 →
Real-world edge suite In progress
The swarm thesis moving toward real deployments — a fleet of small models that learn on-device and share what they learn. Field results in progress.
Note 5 →
For the technical reader
Sources, left to right in the stack. Column 1: 98K/400 KB Sudoku solver at 4,865/4,865 (Note 11); ~37M Hanoi core, 3–8-disk training generalizing to the 30-disk optimum (Notes 3 & 9); ARC program-writer with a 96.7% fresh-transfer verified-program rate (Note 13); 30M→962M ladder at β = 0.245, R² = 0.938 (Note 1). Column 2: 376M brain on the frozen 3B-active Nemotron Nano, GSM8K 25.0→82.5 (a one-week climb of +12.5 → +44.5 → +57.5, completed in Note 14); the same brain line on the 120B Super, 53.5→76.5 (+23, n=200, interim); a 27B host at +26.5 on-device at 4-bit; harm-free across the ARC-E/ARC-C/PIQA/HellaSwag row (Note 13). Column 3: on-device assistant at 42→68% with zero confidently-wrong (Note 6); spaced memory 81% vs 32% at day ten and self-teaching 87→100% (Note 7); portable memory transplant 0→100% and swarm sharing 45.6→100% ×3 (Note 8). Full evidence, configs and logs are available to qualified partners under NDA.