fix: reserve a hybrid forward's full scratch, and report a failed CUDA allocation once #104
Loading…
Reference in a new issue
No description provided.
Delete branch "fix/hybrid-scratch-reserve"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Serving Nemotron 3 Nano with
--max-batch-tokens autocan run the device out of memory on a long prompt, and the failure is reported ascuda: Embedding: out of memory, from an op that allocates nothing. Two separate defects are behind that, one per commit.The scratch reserve is too small for a hybrid
engine.forwardScratchis whatauto:fillandgllm planhold back for a forward step. For a hybrid it reserved the widest of the attention+MLP, Mamba and MoE buffer sets, on the reasoning that layers run one at a time.nemotronh.Forwardallocates every kind's buffers up front for the whole step, so they are all alive together.On Nemotron 3 Nano that is 88,640 F32 values per batched token allocated against 57,984 reserved.
auto:fillgives the KV cache everything the reserve does not claim, so the difference has nowhere to come from. At the default 8192-token budget the gap is under 1 GiB and only a step filled to the budget reaches it;--max-batch-tokens autoraises the budget to as much as the context window, where the gap is tens of GiB and a single long prompt reaches it.The estimate now adds the Mamba and MoE buffers to the attention+MLP term instead of taking the maximum. It stays a slight over-count (the attention+MLP term is the Mistral forward's), which also leaves room for the pooled matmul scratch.
gllm planfor Nemotron 3 Nano, 96 GiB, 16 sequences:--max-model-len 65536 --max-batch-tokens auto: scratch--max-batch-tokens auto: budgetModels without a hybrid pattern are unaffected. An instance that sets neither
--max-model-lennor--kv-cachegets a slightly smaller KV cache (74.9 to 73.7 GiB in the default row above).A failed allocation is reported twice
A failed CUDA runtime call returns its status and also leaves it as the host thread's last error. The only reader is
launch_status(cudaGetLastError), which every op calls after launching its kernel. AcudaMallocthat ran out of memory was therefore reported byAllocand then again by the next op to launch a kernel on that thread, usuallyEmbeddingat the start of the next step, failing a request unrelated to the allocation.cudaErrnow clears the last error as it builds the Go error, andcudaRCgoes through it.Verification
go build ./...,go vet ./...,go test ./...andgofmtare clean.TestForwardScratchCoversHybridForwardruns the tinynemotron_hfixture (all four layer kinds) on a backend that records its allocation high-water mark and requires the reserve to cover it. Against the old estimate it fails with 122,256 bytes allocated and 69,264 reserved.make cudabuilds andgo vet -tags cudatype-checks the cuda-tagged tests.TestCUDAFailedAllocDoesNotFailNextOpreproduces the double report: with the clear removed it fails on an RTX 3070 withEmbedding after a failed Alloc: cuda: Embedding: out of memory, the message seen in serving; with the fix it passes. (The second commit's message predates this run and says the test was only compiled.)./internal/backend/cuda/,./internal/model/...and./internal/engine/on the RTX 3070 (sm_86), and./internal/backend/cuda/and./internal/model/nemotronh/on the RTX PRO 6000 Blackwell (sm_120).🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Nemotron 3 Nano | tracking id
r-bf52ef-626895This is an AI-generated review and may contain mistakes.
Status: ✅ Completed
❓ Verdict: No verdict returned
(No review body returned)