W4M · Research Notes

Research Note · New results · 14 / 14

We made the brain leaner. It scored higher.

Everyone assumes you improve a brain by adding to it. We tried the opposite — took parts out — and braced for a small regression. On the same frozen NVIDIA Nemotron Nano, at the same brain budget, the leaner brain scored the highest number we have measured. One frozen model, one week: +12.5 → +44.5 → +57.5.

W4M Research · July 2026 · ~6 min read

25.0 → 82.5%
GSM8K (flexible extraction, n=200) on the frozen 3B-active Nemotron Nano — before vs after a leaner GoM-Brain. +57.5 points, the biggest lift in the series, a 376M-class brain, not one NVIDIA weight touched.
+12.5 → +44.5 → +57.5
One week, one frozen model, one 376M-class brain budget, one GSM8K axis. The model never grew; the method got sharper (a deeper checkpoint read) and then leaner (this note). Same ground each time — the base reproduces 25.0% to the decimal.
~16% leaner
Fewer moving parts, roughly 16% fewer parameters — and yet the cleanest outputs we have measured (strict-match jumped too) and a clean multiple-choice do-no-harm sweep: every safety row positive on both metrics.
The short version Every instinct in this field pushes one way: to make a component better, add to it. We went the other way. Over the past week we published three numbers on the same frozen NVIDIA Nemotron Nano, at the same 376M-class brain budget, on the same GSM8K axis: +12.5 in Note 13, then +44.5 days later, when a deeper checkpoint-selection instrument re-read the training run. This note is the third turn. We suspected the brain was carrying weight it didn't need, built a leaner variant — fewer moving parts, ~16% fewer parameters — and expected, honestly, to pay a point or two for the simplicity. Instead the slimmer brain scored higher than anything before it: 25.0% → 82.5% on GSM8K (+57.5 points, n=200) — the biggest lift in the series, with the cleanest outputs we have measured and a clean multiple-choice do-no-harm sweep, every safety row positive on both metrics. The frozen base still reproduces our published 25.0% to the decimal, so the before/after is on identical ground. The lesson is the surprising part: the attach mechanism carries the lift, not the brain's bulk. One frozen model, one week: +12.5 → +44.5 → +57.5. The model didn't get bigger — the method got sharper and leaner.

The posture note first, because this series lives or dies on it: every prior number was real and it reproduced. Note 13's +12.5 stands; the +44.5 that followed stands. This is not a walk-back and not a correction — it is the next step in the same direction. What changed is not the model and not the benchmark; it is the brain's design, and it changed in the direction almost nobody bets on: down. Same frozen weights, same 376M-class budget, same GSM8K axis — a leaner brain, read by the same honest instrument.

The climb so far — +12.5, then +44.5

Two notes set the stage. In Note 13 we attached a 376M GoM-Brain to NVIDIA's frozen 3B-active Nemotron Nano and measured +12.5 points on GSM8K — 25.0% → 37.5% — from a brain trained overnight on a single GPU, with zero NVIDIA weights touched. Days later, a deeper checkpoint-selection instrument re-read the same kind of training run on the same frozen model and found nearly three times the lift: +44.5 points, 25.0% → 69.5%, confirmed at n=200. Nothing about Nano changed between those two notes — the way we selected the brain did. This note continues that thread, and it changes something new: the brain itself.

The counterintuitive question — is the brain carrying dead weight?

The reflex in machine learning is additive. To make a component stronger, you give it more — more parameters, more mechanisms, more moving parts. Watching that deep-ladder run, we started to suspect the reverse might be true here: that the brain was carrying subsystems that cost parameters and complexity without earning their keep once the attach mechanism was doing the real work. So we built a leaner variant of the brain — fewer moving parts, roughly 16% fewer parameters. We are deliberately not itemizing what came out; the specific composition is the part we keep. The honest expectation going in was a small regression — you usually trade a little accuracy for that much simplicity. Even a leaner brain that merely held its ground would have been a real win on its own terms: cheaper to train, cheaper to run, easier to fit on Mac- and phone-class hardware. We were braced to give up a point or two to get it.

The leaner brain scored higher — 25.0 → 82.5

It didn't cost a point. It gained the most points we have ever measured on this model. On the same frozen Nano, the same GSM8K axis, and the same single n=200 confirmation that decided the +44.5 result, the leaner brain scored 82.5%+57.5 over the 25.0% base, and a clear step past the 69.5% posted by the full-size brain it replaced.

ConfigurationGSM8K (flexible extraction, n=200)
Frozen Nemotron Nano (30B, 3B-active) — reproduces our published base to the decimal25.0%
+ original GoM-Brain (Note 13's published Nano row)37.5% (+12.5)
+ deep-ladder GoM-Brain (n=200)69.5% (+44.5)
+ leaner GoM-Brain — ~16% fewer parameters82.5% (+57.5)

Two things about that result deserve to be separated. First, it is not merely more often right — it is cleaner. GSM8K can be scored two ways: flexible extraction, which digs the final number out of whatever the model wrote, and strict-match, which demands the answer in exact canonical form. The headline above is flexible extraction — the same conservative axis on which the frozen base reproduces our published 25.0% to the decimal, apples to apples with every prior note. But the leaner brain's strict-match score jumped too, by more than the full-size brain's did: the outputs got tidier, not just luckier. That is the opposite of what you would expect from taking machinery out. Second, the win isn't a GSM8K artifact — the safety rows moved the right way as well.

We re-ran the standard multiple-choice do-no-harm battery in the same environment — ARC-Easy, ARC-Challenge, PIQA, HellaSwag, n=500 each — the check that a math-focused brain hasn't quietly degraded general ability. Every row came back positive, and this time on both metrics at once: raw accuracy and the length-normalized reading a heavier brain had nicked on a row or two. A clean sweep.

Do-no-harm (accuracy, n=500 each)Frozen base+ leaner brain
ARC-Easy75.4%82.2% (+6.8)
ARC-Challenge45.8%55.6% (+9.8)
PIQA81.6%82.0% (+0.4)
HellaSwag51.8%53.4% (+1.6)

That clean-sweep-on-both-metrics detail matters more than it looks. The full-size brain's do-no-harm rows were fine on accuracy but wobbled slightly on the length-normalized reading; the leaner brain erased the wobble — every row positive on accuracy and on the length-normalized metric. Fewer parts, and the failure the extra parts were supposedly insuring against got smaller, not larger.

We took parts out of the brain and it got better — cleaner outputs, a clean safety sweep, the biggest lift in the series. The attach mechanism was carrying the result all along; the bulk was just along for the ride.

Why smaller-and-better is the part that compounds

Line the week up and the shape is the whole point. Same frozen NVIDIA model, same 376M-class brain budget, same benchmark, three notes: +12.5 → +44.5 → +57.5. The model never got bigger. What improved was the method — first how we read the training run, then how the brain itself is built. And this newest step improved it in the direction that matters most for where GoM actually runs: leaner is cheaper to train, cheaper to run, and easier to fit on Mac- and phone-class hardware. The same flywheel turns again — shipping on-device is exactly what makes us ask whether a component is earning its parameters, and this time the answer made the brain both smaller and stronger. The next lever in the same queue is confidence-gated attach — calling the brain only when it helps, which we measured on a sibling base — and a leaner brain makes that lever cheaper still.

For the technical reader Base: NVIDIA Nemotron Nano 30B (3B-active, hybrid Mamba-2 + MoE), public BF16 release, weights frozen throughout. GoM-Brain: a leaner variant of the 376M-class brain studied on this base in Note 13 and the deep-ladder re-read that followed — roughly 16% fewer parameters, attached at one interior layer; we do not disclose which internal components were removed to slim it. Selection used the same deeper instrument that produced the +44.5 read: training carried well past the industry's standard stopping point, dozens of checkpoints evaluated automatically along the way, a battery of independent health checks screening every candidate, and selection rules and accept bars committed to git before the data existed. The selected checkpoint was confirmed at n=200. GSM8K uses flexible answer extraction for the headline, full-forward zero-shot — deliberately conservative, so the frozen base sits below Nano's published chat-mode number; the delta is the measurement, and the frozen base reproduces our published 25.0% to the decimal in this environment, so +57.5 is measured on identical ground. Strict-match improved alongside flexible extraction, by more than the full-size brain's did. Multiple-choice do-no-harm (ARC-Easy, ARC-Challenge, PIQA, HellaSwag, n=500 each) is positive on both accuracy and the length-normalized metric across all four rows. Full configs, eval logs, and the campaign audit ledger are available to qualified partners under NDA.
Honest note Four honesties an informed reader deserves. One — repair, not addition. The brain repairs a weak reasoning path rather than stacking onto a strong one; on a base already scoring high under this harness the same attach measures as a wash (our standing family control). Nano sits squarely in the repair regime, so read +57.5 as headroom recovered, not capability conjured. Two — these are BF16 datacenter numbers, and no edge claim rides on them yet. A datacenter checkpoint can invert at deployed on-device precision — the very lesson that started this instrument — so we make no on-device claim for this leaner Nano brain until we re-run it at deployed precision, even though "leaner" is exactly what helps on-device. Three — the headline is a single n=200 confirmation on the same axis the frozen base reproduces to the decimal. It decides against a bar fixed before the run; the screening scores along the way only narrow the field. Four — this is an instrument-and-design win, not a correction. Every prior number was real and reproduced: Note 13's +12.5 and the +44.5 that followed both stand. A leaner brain read by the same honest instrument is what produced the larger number. NVIDIA had no involvement in this work; the Nemotron name is theirs, the measurements are ours.

Methodology: all figures are from real, logged runs; every claim traces to an entry in the campaign audit ledger. Selection rules and the accept bar were pre-registered before the data existed; the headline is a single n=200 confirmation. GSM8K uses flexible answer extraction for the headline axis (both arms, full-forward, method-matched), with strict-match reported as a cleanliness check; the frozen base reproduces our published 25.0% to the decimal in this environment. Multiple-choice do-no-harm rows are n=500 each and positive on both metrics. The leaner brain carries roughly 16% fewer parameters than the full-size brain; the specific components removed to slim it are not disclosed. No result in this note is extrapolated beyond the configurations tested, and no edge claim is made without a deployed-precision re-evaluation.

Smaller and better, on the same frozen model. The method is the product.

A repeatable, pre-registered onboarding protocol that reads a whole training run honestly — dozens of checkpoints, health checks that reject the ones that lie, accept bars fixed before the data existed — now paired with a leaner brain that trains and runs for less. Point it at any frozen model. Qualified partners can request the full evidence — configs, logs, and the audit ledger.

w4m.ai