feat(engine): reserve hybrid recurrent state in the KV budget and plan #83

Merged
rcsheets merged 1 commit from feat/hybrid-state-sizing into main 2026-09-27 11:39:46 +00:00
Owner

Phase 1 of docs/hybrid-state-cache.md: size the recurrent state that hybrid Mamba2 models carry per sequence, and reserve it before the KV cache gets what is left.

What the state is

A Mamba2 layer keeps a fixed-size state for each running sequence, whatever its length:

  • the SSM state, heads x head_dim x state_size
  • the causal conv's last conv_kernel - 1 inputs, over conv_dim channels

Both are F32, which is what the reference keeps its SSM cache in. On Nemotron 3 Nano that is 2 MiB + 72 KiB per layer across 23 Mamba layers, 47.6 MiB per sequence, about as much as 4,000 tokens of its KV cost.

Changes

  • stateBytesPerSeq(cfg), next to kvBytesPerToken, computes the per-sequence state. It is zero for models without Mamba layers.
  • computeKVBudget reserves the state up front. It takes MaxRunningSeqs x stateBytesPerSeq off the top, next to the fixed forward scratch. auto:fill, --max-batch-tokens auto and gllm plan all go through this function, so they all leave room for it. It is a reserve rather than a share of the remainder because every running sequence holds its state whether or not it has decoded anything.
  • The error names --max-seqs. When the state alone does not fit, the error gives its size and names --max-seqs, the setting that scales it.
  • Config refuses zero-sized state. A pattern with Mamba layers but a missing or zero Mamba dimension (including the new conv_kernel) fails to load. An undercount here is never safe: the missing bytes have already gone to the KV cache, so the first forward OOMs after startup has reported success.
  • gllm plan prints a state cache: line for hybrids, and --json gains state_cache_bytes and state_bytes_per_seq. The per-sequence figure is filled in even when the plan does not fit, so an orchestrator can see how much of the shortfall --max-seqs controls.
  • The design doc marks phase 1 done. Allocating the state, the weights accounting and the memory gauge move to phase 2, where they belong, since that is when the state is actually allocated.

No served model has Mamba layers yet, so behavior is unchanged until a hybrid architecture is registered.

Testing

  • TestStateBytesPerSeq pins Nemotron's 23 x (64·64·128 + 6144·3) x 4 B, and zero for a dense config.
  • TestComputeKVBudgetReservesState checks that the KV cache gets exactly free memory minus scratch, state and margin, and that a state larger than free memory fails with an error naming --max-seqs.
  • TestLoadRejectsMambaWithoutDims covers Mamba layers without conv_kernel (refused), and a pattern without Mamba layers that needs no Mamba dimensions (accepted).
  • go build, go vet and go test ./... pass. No CUDA code changed.

gllm plan can't be run end to end on a hybrid yet, because it measures weights with the model's real loader and nemotron_h isn't registered, so the sizing is tested at the function level.

🤖 Generated with Claude Code

Phase 1 of `docs/hybrid-state-cache.md`: size the recurrent state that hybrid Mamba2 models carry per sequence, and reserve it before the KV cache gets what is left. ## What the state is A Mamba2 layer keeps a fixed-size state for each running sequence, whatever its length: - the SSM state, `heads x head_dim x state_size` - the causal conv's last `conv_kernel - 1` inputs, over `conv_dim` channels Both are F32, which is what the reference keeps its SSM cache in. On Nemotron 3 Nano that is 2 MiB + 72 KiB per layer across 23 Mamba layers, **47.6 MiB per sequence**, about as much as 4,000 tokens of its KV cost. ## Changes - **`stateBytesPerSeq(cfg)`**, next to `kvBytesPerToken`, computes the per-sequence state. It is zero for models without Mamba layers. - **`computeKVBudget` reserves the state up front.** It takes `MaxRunningSeqs x stateBytesPerSeq` off the top, next to the fixed forward scratch. `auto:fill`, `--max-batch-tokens auto` and `gllm plan` all go through this function, so they all leave room for it. It is a reserve rather than a share of the remainder because every running sequence holds its state whether or not it has decoded anything. - **The error names `--max-seqs`.** When the state alone does not fit, the error gives its size and names `--max-seqs`, the setting that scales it. - **Config refuses zero-sized state.** A pattern with Mamba layers but a missing or zero Mamba dimension (including the new `conv_kernel`) fails to load. An undercount here is never safe: the missing bytes have already gone to the KV cache, so the first forward OOMs after startup has reported success. - **`gllm plan`** prints a `state cache:` line for hybrids, and `--json` gains `state_cache_bytes` and `state_bytes_per_seq`. The per-sequence figure is filled in even when the plan does not fit, so an orchestrator can see how much of the shortfall `--max-seqs` controls. - **The design doc** marks phase 1 done. Allocating the state, the weights accounting and the memory gauge move to phase 2, where they belong, since that is when the state is actually allocated. No served model has Mamba layers yet, so behavior is unchanged until a hybrid architecture is registered. ## Testing - `TestStateBytesPerSeq` pins Nemotron's 23 x (64·64·128 + 6144·3) x 4 B, and zero for a dense config. - `TestComputeKVBudgetReservesState` checks that the KV cache gets exactly free memory minus scratch, state and margin, and that a state larger than free memory fails with an error naming `--max-seqs`. - `TestLoadRejectsMambaWithoutDims` covers Mamba layers without `conv_kernel` (refused), and a pattern without Mamba layers that needs no Mamba dimensions (accepted). - `go build`, `go vet` and `go test ./...` pass. No CUDA code changed. `gllm plan` can't be run end to end on a hybrid yet, because it measures weights with the model's real loader and `nemotron_h` isn't registered, so the sizing is tested at the function level. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(engine): reserve hybrid recurrent state in the KV budget and plan
All checks were successful
ci / test_and_build (pull_request) Successful in 23s
0d7d06a2c9
A Mamba2 layer carries a fixed-size recurrent state per running
sequence: its SSM state (heads x head_dim x state_size) plus the causal
conv's last conv_kernel-1 inputs, in F32. On Nemotron 3 Nano that is
47.6 MiB per sequence across 23 layers, regardless of length. This adds
stateBytesPerSeq and takes MaxRunningSeqs of it off the top in
computeKVBudget, next to the fixed forward scratch, so auto:fill and
gllm plan both leave room for it.

It is a reserve rather than part of the leftover because every running
sequence holds its slot whether or not it has decoded anything. That
also means an undercount here is never safe: the missing bytes have
already been given to the KV cache, and the forward OOMs after startup
reports success. So config.Load now refuses a pattern with Mamba layers
if any Mamba dimension (including the new conv_kernel) is missing or
zero, rather than letting the state be sized at zero. When the state
alone doesn't fit, the error says how much it is and names --max-seqs,
since that is the knob that scales it.

gllm plan reports the reserve as its own line, and --json gains
state_cache_bytes and state_bytes_per_seq. No served model has Mamba
layers yet, so this changes nothing until a hybrid architecture is
registered. This is phase 1 of docs/hybrid-state-cache.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b90038-269093
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-b90038-269093.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b90038-269093`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-b90038-269093`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/hybrid-state-sizing 2026-09-27 11:39:46 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!83
No description provided.