The posture note first, because this series lives or dies on it: every prior number was real and it reproduced. Note 13's +12.5 stands; the +44.5 that followed stands. This is not a walk-back and not a correction — it is the next step in the same direction. What changed is not the model and not the benchmark; it is the brain's design, and it changed in the direction almost nobody bets on: down. Same frozen weights, same 376M-class budget, same GSM8K axis — a leaner brain, read by the same honest instrument.
The climb so far — +12.5, then +44.5
Two notes set the stage. In Note 13 we attached a 376M GoM-Brain to NVIDIA's frozen 3B-active Nemotron Nano and measured +12.5 points on GSM8K — 25.0% → 37.5% — from a brain trained overnight on a single GPU, with zero NVIDIA weights touched. Days later, a deeper checkpoint-selection instrument re-read the same kind of training run on the same frozen model and found nearly three times the lift: +44.5 points, 25.0% → 69.5%, confirmed at n=200. Nothing about Nano changed between those two notes — the way we selected the brain did. This note continues that thread, and it changes something new: the brain itself.
The counterintuitive question — is the brain carrying dead weight?
The reflex in machine learning is additive. To make a component stronger, you give it more — more parameters, more mechanisms, more moving parts. Watching that deep-ladder run, we started to suspect the reverse might be true here: that the brain was carrying subsystems that cost parameters and complexity without earning their keep once the attach mechanism was doing the real work. So we built a leaner variant of the brain — fewer moving parts, roughly 16% fewer parameters. We are deliberately not itemizing what came out; the specific composition is the part we keep. The honest expectation going in was a small regression — you usually trade a little accuracy for that much simplicity. Even a leaner brain that merely held its ground would have been a real win on its own terms: cheaper to train, cheaper to run, easier to fit on Mac- and phone-class hardware. We were braced to give up a point or two to get it.
The leaner brain scored higher — 25.0 → 82.5
It didn't cost a point. It gained the most points we have ever measured on this model. On the same frozen Nano, the same GSM8K axis, and the same single n=200 confirmation that decided the +44.5 result, the leaner brain scored 82.5% — +57.5 over the 25.0% base, and a clear step past the 69.5% posted by the full-size brain it replaced.
| Configuration | GSM8K (flexible extraction, n=200) |
|---|---|
| Frozen Nemotron Nano (30B, 3B-active) — reproduces our published base to the decimal | 25.0% |
| + original GoM-Brain (Note 13's published Nano row) | 37.5% (+12.5) |
| + deep-ladder GoM-Brain (n=200) | 69.5% (+44.5) |
| + leaner GoM-Brain — ~16% fewer parameters | 82.5% (+57.5) |
Two things about that result deserve to be separated. First, it is not merely more often right — it is cleaner. GSM8K can be scored two ways: flexible extraction, which digs the final number out of whatever the model wrote, and strict-match, which demands the answer in exact canonical form. The headline above is flexible extraction — the same conservative axis on which the frozen base reproduces our published 25.0% to the decimal, apples to apples with every prior note. But the leaner brain's strict-match score jumped too, by more than the full-size brain's did: the outputs got tidier, not just luckier. That is the opposite of what you would expect from taking machinery out. Second, the win isn't a GSM8K artifact — the safety rows moved the right way as well.
We re-ran the standard multiple-choice do-no-harm battery in the same environment — ARC-Easy, ARC-Challenge, PIQA, HellaSwag, n=500 each — the check that a math-focused brain hasn't quietly degraded general ability. Every row came back positive, and this time on both metrics at once: raw accuracy and the length-normalized reading a heavier brain had nicked on a row or two. A clean sweep.
| Do-no-harm (accuracy, n=500 each) | Frozen base | + leaner brain |
|---|---|---|
| ARC-Easy | 75.4% | 82.2% (+6.8) |
| ARC-Challenge | 45.8% | 55.6% (+9.8) |
| PIQA | 81.6% | 82.0% (+0.4) |
| HellaSwag | 51.8% | 53.4% (+1.6) |
That clean-sweep-on-both-metrics detail matters more than it looks. The full-size brain's do-no-harm rows were fine on accuracy but wobbled slightly on the length-normalized reading; the leaner brain erased the wobble — every row positive on accuracy and on the length-normalized metric. Fewer parts, and the failure the extra parts were supposedly insuring against got smaller, not larger.
Why smaller-and-better is the part that compounds
Line the week up and the shape is the whole point. Same frozen NVIDIA model, same 376M-class brain budget, same benchmark, three notes: +12.5 → +44.5 → +57.5. The model never got bigger. What improved was the method — first how we read the training run, then how the brain itself is built. And this newest step improved it in the direction that matters most for where GoM actually runs: leaner is cheaper to train, cheaper to run, and easier to fit on Mac- and phone-class hardware. The same flywheel turns again — shipping on-device is exactly what makes us ask whether a component is earning its parameters, and this time the answer made the brain both smaller and stronger. The next lever in the same queue is confidence-gated attach — calling the brain only when it helps, which we measured on a sibling base — and a leaner brain makes that lever cheaper still.
Methodology: all figures are from real, logged runs; every claim traces to an entry in the campaign audit ledger. Selection rules and the accept bar were pre-registered before the data existed; the headline is a single n=200 confirmation. GSM8K uses flexible answer extraction for the headline axis (both arms, full-forward, method-matched), with strict-match reported as a cleanliness check; the frozen base reproduces our published 25.0% to the decimal in this environment. Multiple-choice do-no-harm rows are n=500 each and positive on both metrics. The leaner brain carries roughly 16% fewer parameters than the full-size brain; the specific components removed to slim it are not disclosed. No result in this note is extrapolated beyond the configurations tested, and no edge claim is made without a deployed-precision re-evaluation.