feat(engine): give each running sequence of a hybrid model a state slot #84
Loading…
Reference in a new issue
No description provided.
Delete branch "feat/hybrid-state-slots"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Phase 2 of
docs/hybrid-state-cache.md: give every running sequence of a hybrid (Mamba2) model a recurrent state slot, and let the model address it through the batch. Phase 1 (#83) sized and reserved that state; this change hands it out.Slots
kvcache.Config.StateSlotsenables slots,AcquireSlothands one out at admission, andReleasefrees it along with the sequence's blocks. Every exit path (finish, preempt, cancel, admission rollback) already goes throughRelease, so none of them can leak a slot, including any exit path added later.MaxRunningSeqsslots. A slot is held only while its sequence is running, so the pool can't run dry. The state cache adds no admission limit and no reason to preempt. IfAcquireSlotever finds the pool empty, that is a bookkeeping bug: it returns an error, andSchedulereleases the sequence and passes the error up.model.Batch.StateSlotscarries each sequence's slot to the forward pass. It is nil for models without recurrent layers. A slot holds the recurrence over exactly positions[0, firstPos), so a sequence whose first scheduled position is 0 starts from zero state and must not read the slot. Whatever a previous owner left there is never seen, so slots never need zeroing.Engine
model.RecurrentState(AllocStateCache(numSlots)) is the optional interface a hybrid architecture implements.engine.Newrefuses a config with Mamba layers whose architecture doesn't implement it, before any weights load. It then callsAllocStateCache(MaxRunningSeqs)afterAllocKVCache, which is the amountauto:fillreserved in #83.--prefix-cacheis refused for hybrid models, right after the config loads. A prefix hit would reuse the attention layers' KV, but the recurrent state at the end of the prefix was never kept, so the model would write fluent text conditioned on the wrong context and nothing would error. The flag is opt-in, so refusing it is better than silently ignoring it. As a backstop,kvcache.NewManagerforces prefix reuse off whenever state slots are on, so the combination is impossible, not only refused.gllm_memory_state_cache_bytesstate_cache_bytesin/v1/internal/statusstate_cachefield in the perfstats log, for hybrid models onlyDocs
AppendKVis, so a failed step must fail every sequence in it and must never be retried.No registered architecture has Mamba layers yet, so serving behavior is unchanged.
Testing
TestStateSlots: distinct ids; an empty pool or a double acquire is an error;Releasefrees and the slot is reused; releasing an unknown sequence leaves the pool alone.TestStateSlotsDisablePrefixCache: a fully registered prompt still doesn't match when slots are on.TestStateSlotsFollowTheSequence: slots are distinct and stable across decode steps; they are freed on preemption (the recompute takes a fresh one at position 0), on finish, and on cancelling a running sequence; cancelling a waiting sequence leaves the pool alone.TestNoStateSlotsWithoutRecurrentLayers: batches carry no slots when slots are off.TestStateSlotShortageIsAnError: a shortage failsScheduleand leaves the unadmitted sequence holding nothing.TestHybridStateSlotsThroughEngineuses a fake hybrid model that allocates its KV and state from real shapes (per attention layer; per Mamba layer, SSM[slots, heads, headDim, state]and conv[slots, convDim, kernel-1]). It runs five concurrent requests through two slots and checks that slots are distinct within each batch, fixed for each sequence's lifetime, start at position 0, and are all free afterward. Since the fake has no weights, it also checks that KV plus state accounts for every allocated byte.TestPrefixCacheRefusedForHybridandTestRecurrentStateRequiredcover the two startup refusals.go build,go vetandgo test ./...pass, and the new tests are clean under-race. No CUDA code changed.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id
r-b9022c-5e9441This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id
r-b9022c-5e9441.Comment
@pr-reviewer-bot retryto try again.