feat(model): Nemotron-H hybrid (Mamba2 + attention + MoE) on the CPU #86

Merged
rcsheets merged 2 commits from feat/nemotronh-cpu into main 2026-09-27 12:19:47 +00:00
Owner

Phase 4 of docs/hybrid-state-cache.md: internal/model/nemotronh implements NVIDIA's Nemotron-H (NemotronHForCausalLM), a Mamba2 / attention / MoE hybrid, on the CPU backend. It is validated against transformers, and the real Nemotron 3 Nano 30B-A3B NVFP4 checkpoint loads. The PR has two commits.

1. refactor(model): share the sigmoid MoE router

GLM's routing has four steps: sigmoid scores; e_score_correction_bias used only to pick experts; group-limited top-k; the unbiased scores renormalized and scaled. It lived inside GLM's per-token MoE. Nemotron-H routes the same way, so the router moves to internal/model/moe as moe.SigmoidRouter, and GLM now calls it.

  • GLM's goldens still pass.
  • One edge changes: a missing routed_scaling_factor now means 1 (HF's default) where GLM would have multiplied every routed expert by 0. Real GLM configs set it, so their output is unchanged.
  • The new package has its own tests: the bias picks experts without weighting them, weights with and without renormalization, and group limiting.

2. feat(model): Nemotron-H on the CPU

Each layer is an RMSNorm plus one mixer, chosen by the layer pattern:

mixer
M Mamba2: in_proj, then CausalConv1D, SSMScan and GatedRMSNorm, then out_proj. The state lives in the engine's slots (#84), and the ops come from #85
* GQA over the paged KV cache, with no positional encoding, since the Mamba layers carry position
E sigmoid-routed experts plus a shared expert, all non-gated relu² MLPs. Experts no token picked are skipped, which is exact
- one relu² MLP

Supporting changes:

  • quantformat.Loader.NVFP4Rows loads a row range of a packed weight. The range carries its own codes and block scales plus the whole weight's global scale, and the split is exact. The checkpoint's in_proj emits [z | xBC | dt] in one row, and nothing downstream takes a strided view, so the weight is split by output block at load, as mistral4 does for its fused projections. Full-precision in_proj weights, which the real checkpoint has in six excluded layers, are split the same way.
  • config reads transformers 5's layers_block_type list, with both current and legacy names. When transformers re-saves a Nemotron-H config it drops hybrid_override_pattern; without the list, such a config would read as all attention, sizing KV for every layer and running no Mamba layers. If a config carries both and they disagree, it is an error.
  • config also reads the Nemotron-H fields the model needs: layer_norm_epsilon, the shared-expert width, activations and bias flags. New refuses variants it doesn't implement (moe_latent_size, other activations, projection biases) instead of running them wrong.
  • The architecture is registered in cmd/gllm.

Validation

testdata/generate_nemotronh.py builds a tiny hybrid (pattern MEM*E-M, every layer kind) in transformers' NemotronHForCausalLM and writes it out under the real checkpoint's tensor names, as two fixtures:

  • tiny in full precision
  • tiny-nvfp4 in the modelopt NVFP4 dialect the real checkpoint ships. The experts and two Mamba layers' projections are quantized and a third is left full precision, as the real checkpoint excludes some. Its goldens are computed from the dequantized weights.

These all match within 2e-4:

  • one-shot prefill
  • token-by-token decode (every Mamba layer resuming from its slot, every attention layer from the paged cache)
  • prefill in two uneven pieces, the second resuming from state
  • two sequences in one batch on separate slots, one prefilling while the other decodes
  • the NVFP4 fixture, both prefill and decode

Reference quirk: the goldens run transformers' scan as a single chunk. Without mamba_ssm, transformers falls back to a pure-torch chunked scan that is off at the first token of every later chunk. Against a plain sequential recurrence, a single chunk agrees to ~1e-8, while 8-token chunks are ~1e-3 off at exactly positions 8 and 16. gllm's scan is that sequential recurrence, and the real model runs mamba_ssm's kernels, not the fallback. The generator comment and a new AGENTS.md gotcha record this so a later fixture doesn't bake the fallback's error into its goldens.

The generator also passes use_cache=False, because transformers cannot build a cache for a pattern containing an MLP layer (KeyError: 'mlp').

Real checkpoint: gllm plan on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 runs the real loader against the counting backend. It accepts every tensor name and shape and reports 20.0 GiB of resident weights plus 16 × 47.6 MiB of recurrent state. This needs no GPU.

go build, go vet and go test ./... pass. No CUDA code changed.

Not in this PR

  • CUDA: the Mamba2 ops still return ErrNotImplemented on CUDA (phase 5), so serving there fails on the first forward with a 501.
  • Chat: the chat template and reasoning parser are phase 6.
  • Quantization reporting: modelopt checkpoints keep their quantization config in hf_quant_config.json, which nothing reads, so gllm plan and status report none. This is unchanged here and already true of GLM.

🤖 Generated with Claude Code

Phase 4 of `docs/hybrid-state-cache.md`: `internal/model/nemotronh` implements NVIDIA's Nemotron-H (`NemotronHForCausalLM`), a Mamba2 / attention / MoE hybrid, on the CPU backend. It is validated against transformers, and the real Nemotron 3 Nano 30B-A3B NVFP4 checkpoint loads. The PR has two commits. ## 1. `refactor(model)`: share the sigmoid MoE router GLM's routing has four steps: sigmoid scores; `e_score_correction_bias` used only to pick experts; group-limited top-k; the unbiased scores renormalized and scaled. It lived inside GLM's per-token MoE. Nemotron-H routes the same way, so the router moves to `internal/model/moe` as `moe.SigmoidRouter`, and GLM now calls it. - GLM's goldens still pass. - One edge changes: a missing `routed_scaling_factor` now means 1 (HF's default) where GLM would have multiplied every routed expert by 0. Real GLM configs set it, so their output is unchanged. - The new package has its own tests: the bias picks experts without weighting them, weights with and without renormalization, and group limiting. ## 2. `feat(model)`: Nemotron-H on the CPU Each layer is an RMSNorm plus one mixer, chosen by the layer pattern: | | mixer | |---|---| | `M` | Mamba2: `in_proj`, then `CausalConv1D`, `SSMScan` and `GatedRMSNorm`, then `out_proj`. The state lives in the engine's slots (#84), and the ops come from #85 | | `*` | GQA over the paged KV cache, with **no positional encoding**, since the Mamba layers carry position | | `E` | sigmoid-routed experts plus a shared expert, all non-gated relu² MLPs. Experts no token picked are skipped, which is exact | | `-` | one relu² MLP | Supporting changes: - **`quantformat.Loader.NVFP4Rows`** loads a row range of a packed weight. The range carries its own codes and block scales plus the whole weight's global scale, and the split is exact. The checkpoint's `in_proj` emits `[z | xBC | dt]` in one row, and nothing downstream takes a strided view, so the weight is split by output block at load, as `mistral4` does for its fused projections. Full-precision `in_proj` weights, which the real checkpoint has in six excluded layers, are split the same way. - **`config` reads transformers 5's `layers_block_type` list**, with both current and legacy names. When transformers re-saves a Nemotron-H config it drops `hybrid_override_pattern`; without the list, such a config would read as all attention, sizing KV for every layer and running no Mamba layers. If a config carries both and they disagree, it is an error. - **`config` also reads the Nemotron-H fields the model needs**: `layer_norm_epsilon`, the shared-expert width, activations and bias flags. `New` refuses variants it doesn't implement (`moe_latent_size`, other activations, projection biases) instead of running them wrong. - **The architecture is registered** in `cmd/gllm`. ## Validation `testdata/generate_nemotronh.py` builds a tiny hybrid (pattern `MEM*E-M`, every layer kind) in transformers' `NemotronHForCausalLM` and writes it out under the real checkpoint's tensor names, as two fixtures: - **`tiny`** in full precision - **`tiny-nvfp4`** in the modelopt NVFP4 dialect the real checkpoint ships. The experts and two Mamba layers' projections are quantized and a third is left full precision, as the real checkpoint excludes some. Its goldens are computed from the dequantized weights. These all match within 2e-4: - one-shot prefill - token-by-token decode (every Mamba layer resuming from its slot, every attention layer from the paged cache) - prefill in two uneven pieces, the second resuming from state - two sequences in one batch on separate slots, one prefilling while the other decodes - the NVFP4 fixture, both prefill and decode **Reference quirk: the goldens run transformers' scan as a single chunk.** Without `mamba_ssm`, transformers falls back to a pure-torch chunked scan that is off at the first token of every later chunk. Against a plain sequential recurrence, a single chunk agrees to ~1e-8, while 8-token chunks are ~1e-3 off at exactly positions 8 and 16. gllm's scan is that sequential recurrence, and the real model runs `mamba_ssm`'s kernels, not the fallback. The generator comment and a new AGENTS.md gotcha record this so a later fixture doesn't bake the fallback's error into its goldens. The generator also passes `use_cache=False`, because transformers cannot build a cache for a pattern containing an MLP layer (`KeyError: 'mlp'`). **Real checkpoint:** `gllm plan` on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 runs the real loader against the counting backend. It accepts every tensor name and shape and reports 20.0 GiB of resident weights plus 16 × 47.6 MiB of recurrent state. This needs no GPU. `go build`, `go vet` and `go test ./...` pass. No CUDA code changed. ## Not in this PR - **CUDA:** the Mamba2 ops still return `ErrNotImplemented` on CUDA (phase 5), so serving there fails on the first forward with a 501. - **Chat:** the chat template and reasoning parser are phase 6. - **Quantization reporting:** modelopt checkpoints keep their quantization config in `hf_quant_config.json`, which nothing reads, so `gllm plan` and status report `none`. This is unchanged here and already true of GLM. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
GLM's routing was sigmoid scores, the e_score_correction_bias used only
to choose experts, group-limited top-k, then the unbiased scores
renormalized and scaled. It lived inside GLM's per-token MoE as methods
on the model. Nemotron-H routes exactly the same way, so it moves to
internal/model/moe as moe.SigmoidRouter, which GLM now calls.

One edge changes: a missing routed_scaling_factor now means 1 (HF's
default), where GLM would have multiplied every routed expert by 0.
Real GLM configs set it (2.5), so their output is unchanged, and the GLM
goldens still pass.

The package has its own tests: the bias picks experts without
weighting them, weights with and without renormalization, and group
limiting keeping the best expert out when its group ranks lower.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
feat(model): Nemotron-H hybrid (Mamba2 + attention + MoE) on the CPU
All checks were successful
ci / test_and_build (pull_request) Successful in 25s
743d07d9dc
Phase 4 of docs/hybrid-state-cache.md: internal/model/nemotronh
implements NemotronHForCausalLM. Each layer is one RMSNorm and one
mixer, chosen by the layer pattern:

- 'M' Mamba2: in_proj, then the causal conv, the selective scan and
  the gated norm, then out_proj. The state lives in the engine's slots.
- '*' attention: plain GQA over the paged KV cache, with no positional
  encoding (the Mamba layers carry position).
- 'E' MoE: sigmoid-routed experts plus a shared expert, all non-gated
  relu^2 MLPs. Experts no token picked are skipped, which is exact.
- '-' MLP: one relu^2 MLP.

Supporting changes:

- quantformat.Loader.NVFP4Rows loads a row range of a packed weight,
  with its own codes and block scales and the whole weight's global.
  in_proj emits [z | xBC | dt] in one row, and nothing downstream takes
  a strided view, so it is split by output block at load, as mistral4
  does for its fused projections.
- config reads transformers 5's layers_block_type list (current and
  legacy names). A re-saved Nemotron-H config drops
  hybrid_override_pattern, and without the list it would read as all
  attention. It also reads the Nemotron-H fields the model needs:
  layer_norm_epsilon, the shared-expert width, activations and bias
  flags. New() refuses variants it doesn't implement (moe_latent_size,
  other activations, projection biases) instead of running them wrong.

Validation is against transformers' NemotronHForCausalLM on two tiny
fixtures with every layer kind: one in full precision, and one in the
modelopt NVFP4 dialect the real checkpoint ships (experts and two
Mamba layers' projections quantized, a third left full precision, as
the real checkpoint excludes some). Prefill, token-by-token decode,
prefill in two pieces, and two sequences batched on separate slots
(one prefilling while the other decodes) all match within 2e-4.

The goldens run the reference scan as one chunk. transformers'
pure-torch chunked scan, its fallback without mamba_ssm, is off by
about 1e-3 at the first token of every later chunk. Checked against a
plain sequential recurrence, a single chunk agrees to ~1e-8, and 8-token
chunks are off at exactly positions 8 and 16. gllm's scan is that
sequential recurrence, and the real model runs mamba_ssm's kernels, not
the fallback. The generator and AGENTS.md record this so a later
fixture doesn't bake the fallback's error into its goldens. The
generator also has to pass use_cache=False, because transformers
can't build a cache for a pattern with an MLP layer.

The real Nemotron 3 Nano 30B-A3B NVFP4 checkpoint loads under gllm plan
(every tensor name and shape accepted, 20.0 GiB of weights, 16 x 47.6
MiB of state). Serving it on CUDA still stops at the Mamba2 ops, which
return ErrNotImplemented until phase 5, and chat needs its template
(phase 6).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b9096f-e4650b
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-b9096f-e4650b.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b9096f-e4650b`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-b9096f-e4650b`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/nemotronh-cpu 2026-09-27 12:19:48 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!86
No description provided.