feat(engine): give each running sequence of a hybrid model a state slot #84

Merged
rcsheets merged 1 commit from feat/hybrid-state-slots into main 2026-09-27 11:56:16 +00:00
Owner

Phase 2 of docs/hybrid-state-cache.md: give every running sequence of a hybrid (Mamba2) model a recurrent state slot, and let the model address it through the batch. Phase 1 (#83) sized and reserved that state; this change hands it out.

Slots

  • The kvcache Manager hands them out. kvcache.Config.StateSlots enables slots, AcquireSlot hands one out at admission, and Release frees it along with the sequence's blocks. Every exit path (finish, preempt, cancel, admission rollback) already goes through Release, so none of them can leak a slot, including any exit path added later.
  • The pool is MaxRunningSeqs slots. A slot is held only while its sequence is running, so the pool can't run dry. The state cache adds no admission limit and no reason to preempt. If AcquireSlot ever finds the pool empty, that is a bookkeeping bug: it returns an error, and Schedule releases the sequence and passes the error up.
  • model.Batch.StateSlots carries each sequence's slot to the forward pass. It is nil for models without recurrent layers. A slot holds the recurrence over exactly positions [0, firstPos), so a sequence whose first scheduled position is 0 starts from zero state and must not read the slot. Whatever a previous owner left there is never seen, so slots never need zeroing.

Engine

  • model.RecurrentState (AllocStateCache(numSlots)) is the optional interface a hybrid architecture implements. engine.New refuses a config with Mamba layers whose architecture doesn't implement it, before any weights load. It then calls AllocStateCache(MaxRunningSeqs) after AllocKVCache, which is the amount auto:fill reserved in #83.
  • --prefix-cache is refused for hybrid models, right after the config loads. A prefix hit would reuse the attention layers' KV, but the recurrent state at the end of the prefix was never kept, so the model would write fluent text conditioned on the wrong context and nothing would error. The flag is opt-in, so refusing it is better than silently ignoring it. As a backstop, kvcache.NewManager forces prefix reuse off whenever state slots are on, so the combination is impossible, not only refused.
  • Memory accounting: weights are now "alloc after load minus KV minus state", and scratch subtracts the state too. The state appears as:
    • the gauge gllm_memory_state_cache_bytes
    • state_cache_bytes in /v1/internal/status
    • a state_cache field in the perfstats log, for hybrid models only

Docs

  • AGENTS.md describes the slots, the prefix-cache exclusion, the new gauge, and a rule: a recurrent update is not idempotent the way AppendKV is, so a failed step must fail every sequence in it and must never be retried.
  • The design doc marks phases 0-2 done.

No registered architecture has Mamba layers yet, so serving behavior is unchanged.

Testing

  • kvcache:
    • TestStateSlots: distinct ids; an empty pool or a double acquire is an error; Release frees and the slot is reused; releasing an unknown sequence leaves the pool alone.
    • TestStateSlotsDisablePrefixCache: a fully registered prompt still doesn't match when slots are on.
  • scheduler:
    • TestStateSlotsFollowTheSequence: slots are distinct and stable across decode steps; they are freed on preemption (the recompute takes a fresh one at position 0), on finish, and on cancelling a running sequence; cancelling a waiting sequence leaves the pool alone.
    • TestNoStateSlotsWithoutRecurrentLayers: batches carry no slots when slots are off.
    • TestStateSlotShortageIsAnError: a shortage fails Schedule and leaves the unadmitted sequence holding nothing.
  • engine:
    • TestHybridStateSlotsThroughEngine uses a fake hybrid model that allocates its KV and state from real shapes (per attention layer; per Mamba layer, SSM [slots, heads, headDim, state] and conv [slots, convDim, kernel-1]). It runs five concurrent requests through two slots and checks that slots are distinct within each batch, fixed for each sequence's lifetime, start at position 0, and are all free afterward. Since the fake has no weights, it also checks that KV plus state accounts for every allocated byte.
    • TestPrefixCacheRefusedForHybrid and TestRecurrentStateRequired cover the two startup refusals.
  • go build, go vet and go test ./... pass, and the new tests are clean under -race. No CUDA code changed.

🤖 Generated with Claude Code

Phase 2 of `docs/hybrid-state-cache.md`: give every running sequence of a hybrid (Mamba2) model a recurrent state slot, and let the model address it through the batch. Phase 1 (#83) sized and reserved that state; this change hands it out. ## Slots - **The kvcache Manager hands them out.** `kvcache.Config.StateSlots` enables slots, `AcquireSlot` hands one out at admission, and `Release` frees it along with the sequence's blocks. Every exit path (finish, preempt, cancel, admission rollback) already goes through `Release`, so none of them can leak a slot, including any exit path added later. - **The pool is `MaxRunningSeqs` slots.** A slot is held only while its sequence is running, so the pool can't run dry. The state cache adds no admission limit and no reason to preempt. If `AcquireSlot` ever finds the pool empty, that is a bookkeeping bug: it returns an error, and `Schedule` releases the sequence and passes the error up. - **`model.Batch.StateSlots`** carries each sequence's slot to the forward pass. It is nil for models without recurrent layers. A slot holds the recurrence over exactly positions `[0, firstPos)`, so a sequence whose first scheduled position is 0 starts from zero state and must not read the slot. Whatever a previous owner left there is never seen, so slots never need zeroing. ## Engine - **`model.RecurrentState`** (`AllocStateCache(numSlots)`) is the optional interface a hybrid architecture implements. `engine.New` refuses a config with Mamba layers whose architecture doesn't implement it, before any weights load. It then calls `AllocStateCache(MaxRunningSeqs)` after `AllocKVCache`, which is the amount `auto:fill` reserved in #83. - **`--prefix-cache` is refused for hybrid models**, right after the config loads. A prefix hit would reuse the attention layers' KV, but the recurrent state at the end of the prefix was never kept, so the model would write fluent text conditioned on the wrong context and nothing would error. The flag is opt-in, so refusing it is better than silently ignoring it. As a backstop, `kvcache.NewManager` forces prefix reuse off whenever state slots are on, so the combination is impossible, not only refused. - **Memory accounting**: weights are now "alloc after load minus KV minus state", and scratch subtracts the state too. The state appears as: - the gauge `gllm_memory_state_cache_bytes` - `state_cache_bytes` in `/v1/internal/status` - a `state_cache` field in the perfstats log, for hybrid models only ## Docs - AGENTS.md describes the slots, the prefix-cache exclusion, the new gauge, and a rule: a recurrent update is not idempotent the way `AppendKV` is, so a failed step must fail every sequence in it and must never be retried. - The design doc marks phases 0-2 done. No registered architecture has Mamba layers yet, so serving behavior is unchanged. ## Testing - **kvcache**: - `TestStateSlots`: distinct ids; an empty pool or a double acquire is an error; `Release` frees and the slot is reused; releasing an unknown sequence leaves the pool alone. - `TestStateSlotsDisablePrefixCache`: a fully registered prompt still doesn't match when slots are on. - **scheduler**: - `TestStateSlotsFollowTheSequence`: slots are distinct and stable across decode steps; they are freed on preemption (the recompute takes a fresh one at position 0), on finish, and on cancelling a running sequence; cancelling a waiting sequence leaves the pool alone. - `TestNoStateSlotsWithoutRecurrentLayers`: batches carry no slots when slots are off. - `TestStateSlotShortageIsAnError`: a shortage fails `Schedule` and leaves the unadmitted sequence holding nothing. - **engine**: - `TestHybridStateSlotsThroughEngine` uses a fake hybrid model that allocates its KV and state from real shapes (per attention layer; per Mamba layer, SSM `[slots, heads, headDim, state]` and conv `[slots, convDim, kernel-1]`). It runs five concurrent requests through two slots and checks that slots are distinct within each batch, fixed for each sequence's lifetime, start at position 0, and are all free afterward. Since the fake has no weights, it also checks that KV plus state accounts for every allocated byte. - `TestPrefixCacheRefusedForHybrid` and `TestRecurrentStateRequired` cover the two startup refusals. - `go build`, `go vet` and `go test ./...` pass, and the new tests are clean under `-race`. No CUDA code changed. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
feat(engine): give each running sequence of a hybrid model a state slot
All checks were successful
ci / test_and_build (pull_request) Successful in 25s
942ec1131f
Phase 2 of docs/hybrid-state-cache.md. A Mamba2 layer keeps fixed-size
recurrent state per running sequence, so the engine now hands each one
a slot and lets the model address it through the batch:

- kvcache.Config.StateSlots makes the Manager hand out slots with
  AcquireSlot, and Release frees a sequence's slot along with its
  blocks. Every exit path (finish, preempt, cancel, admission
  rollback) already goes through Release, so none of them can leak a
  slot, now or later.
- The engine sets StateSlots to MaxRunningSeqs. A slot is held only
  while a sequence runs, so the pool can't run dry: running out is a
  bookkeeping bug, which AcquireSlot reports and Schedule passes up
  after releasing the sequence. It is never a reason to preempt.
- model.Batch.StateSlots carries each sequence's slot to the forward
  pass. A slot holds the recurrence over exactly positions
  [0, firstPos), so a sequence starting at position 0 must not read
  its slot, and whatever a previous owner left there is never seen.
- model.RecurrentState (AllocStateCache) is the optional interface a
  hybrid architecture implements. engine.New refuses a config with
  Mamba layers whose architecture doesn't implement it, before any
  weights load, and calls it after AllocKVCache with MaxRunningSeqs
  slots, which is what auto:fill reserved.
- --prefix-cache is refused for hybrid models, right after the config
  loads. A prefix hit would reuse KV with no recurrent state behind
  it, so the model would write fluent text conditioned on the wrong
  context and nothing would error. As a backstop, NewManager forces
  prefix reuse off whenever state slots are on.
- Memory accounting subtracts the state from weights and scratch, and
  exposes it as gllm_memory_state_cache_bytes, state_cache_bytes in
  /v1/internal/status, and a state_cache field in the perfstats log
  (hybrids only).

No registered architecture has Mamba layers yet, so serving behavior is
unchanged. Tests cover the Manager's slot lifecycle and prefix
backstop, every scheduler exit path, a shortage surfacing as an
error, and an end-to-end engine run over a fake hybrid. The fake
allocates its caches from shapes, so the byte accounting is checked
against real allocations. It runs five requests through two slots and
checks that slots are distinct within a batch, stable for a sequence's
lifetime, start at position 0, and all come back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b9022c-5e9441
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-b9022c-5e9441.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b9022c-5e9441`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-b9022c-5e9441`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch feat/hybrid-state-slots 2026-09-27 11:56:16 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!84
No description provided.