feat(engine): reserve hybrid recurrent state in the KV budget and plan #83
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/hybrid-state-sizing"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Phase 1 of
docs/hybrid-state-cache.md: size the recurrent state that hybrid Mamba2 models carry per sequence, and reserve it before the KV cache gets what is left.What the state is
A Mamba2 layer keeps a fixed-size state for each running sequence, whatever its length:
heads x head_dim x state_sizeconv_kernel - 1inputs, overconv_dimchannelsBoth are F32, which is what the reference keeps its SSM cache in. On Nemotron 3 Nano that is 2 MiB + 72 KiB per layer across 23 Mamba layers, 47.6 MiB per sequence, about as much as 4,000 tokens of its KV cost.
Changes
stateBytesPerSeq(cfg), next tokvBytesPerToken, computes the per-sequence state. It is zero for models without Mamba layers.computeKVBudgetreserves the state up front. It takesMaxRunningSeqs x stateBytesPerSeqoff the top, next to the fixed forward scratch.auto:fill,--max-batch-tokens autoandgllm planall go through this function, so they all leave room for it. It is a reserve rather than a share of the remainder because every running sequence holds its state whether or not it has decoded anything.--max-seqs. When the state alone does not fit, the error gives its size and names--max-seqs, the setting that scales it.conv_kernel) fails to load. An undercount here is never safe: the missing bytes have already gone to the KV cache, so the first forward OOMs after startup has reported success.gllm planprints astate cache:line for hybrids, and--jsongainsstate_cache_bytesandstate_bytes_per_seq. The per-sequence figure is filled in even when the plan does not fit, so an orchestrator can see how much of the shortfall--max-seqscontrols.No served model has Mamba layers yet, so behavior is unchanged until a hybrid architecture is registered.
Testing
TestStateBytesPerSeqpins Nemotron's 23 x (64·64·128 + 6144·3) x 4 B, and zero for a dense config.TestComputeKVBudgetReservesStatechecks that the KV cache gets exactly free memory minus scratch, state and margin, and that a state larger than free memory fails with an error naming--max-seqs.TestLoadRejectsMambaWithoutDimscovers Mamba layers withoutconv_kernel(refused), and a pattern without Mamba layers that needs no Mamba dimensions (accepted).go build,go vetandgo test ./...pass. No CUDA code changed.gllm plancan't be run end to end on a hybrid yet, because it measures weights with the model's real loader andnemotron_hisn't registered, so the sizing is tested at the function level.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id
r-b90038-269093This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id
r-b90038-269093.Comment
@pr-reviewer-bot retryto try again.