fix: reserve a hybrid forward's full scratch, and report a failed CUDA allocation once #104

Merged
rcsheets merged 2 commits from fix/hybrid-scratch-reserve into main 2026-10-02 07:27:17 +00:00
Owner

Serving Nemotron 3 Nano with --max-batch-tokens auto can run the device out of memory on a long prompt, and the failure is reported as cuda: Embedding: out of memory, from an op that allocates nothing. Two separate defects are behind that, one per commit.

The scratch reserve is too small for a hybrid

engine.forwardScratch is what auto:fill and gllm plan hold back for a forward step. For a hybrid it reserved the widest of the attention+MLP, Mamba and MoE buffer sets, on the reasoning that layers run one at a time. nemotronh.Forward allocates every kind's buffers up front for the whole step, so they are all alive together.

On Nemotron 3 Nano that is 88,640 F32 values per batched token allocated against 57,984 reserved. auto:fill gives the KV cache everything the reserve does not claim, so the difference has nowhere to come from. At the default 8192-token budget the gap is under 1 GiB and only a step filled to the budget reaches it; --max-batch-tokens auto raises the budget to as much as the context window, where the gap is tens of GiB and a single long prompt reaches it.

The estimate now adds the Mamba and MoE buffers to the attention+MLP term instead of taking the maximum. It stays a slight over-count (the attention+MLP term is the Mistral forward's), which also leaves room for the pooled matmul scratch.

gllm plan for Nemotron 3 Nano, 96 GiB, 16 sequences:

Setting Before After
default budget (8192): scratch 1.8 GiB 3.0 GiB
--max-model-len 65536 --max-batch-tokens auto: scratch 14.7 GiB 24.2 GiB
--max-batch-tokens auto: budget 171,560 tokens 104,090 tokens

Models without a hybrid pattern are unaffected. An instance that sets neither --max-model-len nor --kv-cache gets a slightly smaller KV cache (74.9 to 73.7 GiB in the default row above).

A failed allocation is reported twice

A failed CUDA runtime call returns its status and also leaves it as the host thread's last error. The only reader is launch_status (cudaGetLastError), which every op calls after launching its kernel. A cudaMalloc that ran out of memory was therefore reported by Alloc and then again by the next op to launch a kernel on that thread, usually Embedding at the start of the next step, failing a request unrelated to the allocation.

cudaErr now clears the last error as it builds the Go error, and cudaRC goes through it.

Verification

  • go build ./..., go vet ./..., go test ./... and gofmt are clean.
  • TestForwardScratchCoversHybridForward runs the tiny nemotron_h fixture (all four layer kinds) on a backend that records its allocation high-water mark and requires the reserve to cover it. Against the old estimate it fails with 122,256 bytes allocated and 69,264 reserved.
  • make cuda builds and go vet -tags cuda type-checks the cuda-tagged tests.
  • TestCUDAFailedAllocDoesNotFailNextOp reproduces the double report: with the clear removed it fails on an RTX 3070 with Embedding after a failed Alloc: cuda: Embedding: out of memory, the message seen in serving; with the fix it passes. (The second commit's message predates this run and says the test was only compiled.)
  • The cuda-tagged tests pass on both cards: ./internal/backend/cuda/, ./internal/model/... and ./internal/engine/ on the RTX 3070 (sm_86), and ./internal/backend/cuda/ and ./internal/model/nemotronh/ on the RTX PRO 6000 Blackwell (sm_120).
  • The real peak of a 40k+ token Nemotron prefill has not been measured on a device either; the reserve is checked against the forward's own allocations, not against device memory, and issue #89 (unaccounted GPU memory) is separate.

🤖 Generated with Claude Code

Serving Nemotron 3 Nano with `--max-batch-tokens auto` can run the device out of memory on a long prompt, and the failure is reported as `cuda: Embedding: out of memory`, from an op that allocates nothing. Two separate defects are behind that, one per commit. ## The scratch reserve is too small for a hybrid `engine.forwardScratch` is what `auto:fill` and `gllm plan` hold back for a forward step. For a hybrid it reserved the widest of the attention+MLP, Mamba and MoE buffer sets, on the reasoning that layers run one at a time. `nemotronh.Forward` allocates every kind's buffers up front for the whole step, so they are all alive together. On Nemotron 3 Nano that is 88,640 F32 values per batched token allocated against 57,984 reserved. `auto:fill` gives the KV cache everything the reserve does not claim, so the difference has nowhere to come from. At the default 8192-token budget the gap is under 1 GiB and only a step filled to the budget reaches it; `--max-batch-tokens auto` raises the budget to as much as the context window, where the gap is tens of GiB and a single long prompt reaches it. The estimate now adds the Mamba and MoE buffers to the attention+MLP term instead of taking the maximum. It stays a slight over-count (the attention+MLP term is the Mistral forward's), which also leaves room for the pooled matmul scratch. `gllm plan` for Nemotron 3 Nano, 96 GiB, 16 sequences: | Setting | Before | After | |---|---|---| | default budget (8192): scratch | 1.8 GiB | 3.0 GiB | | `--max-model-len 65536 --max-batch-tokens auto`: scratch | 14.7 GiB | 24.2 GiB | | `--max-batch-tokens auto`: budget | 171,560 tokens | 104,090 tokens | Models without a hybrid pattern are unaffected. An instance that sets neither `--max-model-len` nor `--kv-cache` gets a slightly smaller KV cache (74.9 to 73.7 GiB in the default row above). ## A failed allocation is reported twice A failed CUDA runtime call returns its status and also leaves it as the host thread's last error. The only reader is `launch_status` (`cudaGetLastError`), which every op calls after launching its kernel. A `cudaMalloc` that ran out of memory was therefore reported by `Alloc` and then again by the next op to launch a kernel on that thread, usually `Embedding` at the start of the next step, failing a request unrelated to the allocation. `cudaErr` now clears the last error as it builds the Go error, and `cudaRC` goes through it. ## Verification - `go build ./...`, `go vet ./...`, `go test ./...` and `gofmt` are clean. - `TestForwardScratchCoversHybridForward` runs the tiny `nemotron_h` fixture (all four layer kinds) on a backend that records its allocation high-water mark and requires the reserve to cover it. Against the old estimate it fails with 122,256 bytes allocated and 69,264 reserved. - `make cuda` builds and `go vet -tags cuda` type-checks the cuda-tagged tests. - `TestCUDAFailedAllocDoesNotFailNextOp` reproduces the double report: with the clear removed it fails on an RTX 3070 with `Embedding after a failed Alloc: cuda: Embedding: out of memory`, the message seen in serving; with the fix it passes. (The second commit's message predates this run and says the test was only compiled.) - The cuda-tagged tests pass on both cards: `./internal/backend/cuda/`, `./internal/model/...` and `./internal/engine/` on the RTX 3070 (sm_86), and `./internal/backend/cuda/` and `./internal/model/nemotronh/` on the RTX PRO 6000 Blackwell (sm_120). - The real peak of a 40k+ token Nemotron prefill has not been measured on a device either; the reserve is checked against the forward's own allocations, not against device memory, and issue #89 (unaccounted GPU memory) is separate. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
forwardScratch reserved the widest of a hybrid's attention+MLP, Mamba and
MoE buffer sets, on the reasoning that layers run one at a time. But
nemotronh.Forward allocates every kind's buffers up front for the whole
step, so they are all alive together: on Nemotron 3 Nano it allocates
88,640 F32 values per batched token against the 57,984 reserved.

auto:fill hands the KV cache whatever the reserve does not claim, so the
shortfall had nowhere to come from, and a prefill near the batch budget
ran the device out of memory. --max-batch-tokens auto made that reachable
by raising the budget from 8192 to as much as the context window.

Add the Mamba and MoE buffers to the attention+MLP estimate instead of
taking the maximum. For Nemotron 3 Nano on 96 GiB with 16 sequences,
gllm plan now reports:

  default budget (8192)        scratch 1.8 -> 3.0 GiB
  auto, --max-model-len 65536  scratch 14.7 -> 24.2 GiB
  auto, no --max-model-len     budget 171560 -> 104090 tokens

Models without a hybrid pattern are unaffected.

TestForwardScratchCoversHybridForward runs the tiny nemotron_h fixture,
which has all four layer kinds, on a backend that records its allocation
high-water mark, and requires the reserve to cover it. Against the old
estimate it fails: 122,256 bytes allocated, 69,264 reserved.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(cuda): clear the runtime's last error when a failure is reported
All checks were successful
ci / test_and_build (pull_request) Successful in 2m4s
94e5acb0ae
A failed CUDA runtime call returns its status and also leaves it as the
host thread's last error, which stays set until something reads it. The
only reader is launch_status (cudaGetLastError), which every op calls
after launching its kernel. So a cudaMalloc that ran out of memory was
reported once by Alloc and then a second time by the next op to launch a
kernel on that thread -- typically Embedding, the first op of the next
step, which surfaced as "cuda: Embedding: out of memory" on a request
unrelated to the allocation that failed, naming an op that allocates
nothing.

cudaErr now clears the last error as it builds the Go error, and cudaRC
goes through it, so each failure is reported by the call that hit it and
only by that call.

TestCUDAFailedAllocDoesNotFailNextOp asks for an allocation the device
cannot satisfy and then requires an Embedding on the same thread to
succeed. It is cuda-tagged and has been compiled but not run on a GPU.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Nemotron 3 Nano | tracking id r-bf52ef-626895
This is an AI-generated review and may contain mistakes.

Status: ✅ Completed


❓ Verdict: No verdict returned

(No review body returned)

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Nemotron 3 Nano | tracking id `r-bf52ef-626895`* *This is an AI-generated review and may contain mistakes.* **Status:** ✅ Completed --- **❓ Verdict: No verdict returned** (No review body returned)
rcsheets deleted branch fix/hybrid-scratch-reserve 2026-10-02 07:27:18 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!104
No description provided.