perf: cut host overhead from long prefills (routing, matmul scratch) #92

Merged
rcsheets merged 1 commit from perf/host-overhead into main 2026-09-27 19:08:25 +00:00
Owner

Cuts the host-side overhead out of long prefills. A 4314-token Nemotron 3 Nano prefill on the RTX PRO 6000 goes from 2.39 s to 1.40 s, and the GPU is now busy 86% of the time, up from about 42%.

What the profile showed

After #91, the prefill's Nsight trace showed the GPU busy for only about 1.0 s of 2.39 s. Splitting its idle time by what the host was doing showed two costs, together about 1.2 s:

per prefill idle time cause
MoE routing ~0.75 s After each MoE layer's router logits came back to the host, the GPU waited a mean 32 ms while every token was routed one at a time. Each token did a full stable sort of all 128 experts and six heap allocations: 5.4 µs per token.
matmul scratch ~0.47 s Each large NVFP4 or bf16 matmul cudaMalloc'd and cudaFree'd a bf16 activation copy and an unpacked weight: about 10,600 pairs per prefill, and cudaFree synchronizes.
launch/sync, other host ~0.15 s about 12,000 ops, each synchronizing before the next launch (not addressed here)

Changes

Routing.

  • moe.SigmoidRouter.RouteInto picks the top k directly: linear in the expert count, ties still broken toward the smaller index. It writes into caller-owned space and allocates nothing. Route wraps it, so GLM is unchanged. 0.79 µs per token, down from 5.4 µs.
  • nemotronh routes a step's tokens across goroutines once there are 256 or more (a decode step stays on the calling goroutine). The packing then reads the picks back instead of routing inline.

Scratch pool (CUDA).

  • The backend keeps a grow-only pool with one buffer per purpose: the bf16 activation, the unpacked bf16 or f32 weight, and W8A8's quantized activation and its scales. The matmul paths take views of these instead of allocating. Each buffer grows by at least half again when it grows, so rising sizes don't reallocate on every call.
  • Reuse is safe because every op finishes on the device before it returns.
  • The buffers still count as live allocations for backend.MemStats, as the transient scratch did before, and Close releases them.

Results

The same 4314-token prefill, measured with the same Nsight setup as #91:

before after
wall time 2.39 s 1.40 s
GPU idle ~1.45 s 0.19 s
cudaMalloc/cudaFree pairs ~10,700 164
routing stall per MoE layer ~32 ms ~2.3 ms
GPU busy share of wall time ~42% 86%
  • Output: unchanged, and chat still answers correctly.
  • Now GPU-bound. The remaining time is kernels: paged attention 0.47 s (the decode flash kernel on a 4k prompt), the NVFP4 unpack 0.22 s, the NVFP4 GEMV 0.17 s, scan and conv 0.14 s.
  • Decode is unchanged at about 23 ms/token. Its cost lies elsewhere, most likely each op's synchronization. It has not been profiled yet.

Testing

  • moe:
    • TestRouteIntoMatchesReference checks RouteInto against a full-sort reference over 8,000 tokens. The logits are coarse so ties are common; the cases also include group masking that drives experts to −inf, and a k larger than the expert count.
    • TestRouteIntoDoesNotAllocate checks that routing a token allocates nothing.
  • nemotronh: TestRouteParallelMatchesSerial routes 1000 tokens through the parallel path and compares each with Route; it passes under -race.
  • cuda: TestCUDAScratchReuse runs NVFP4 matmuls whose sizes rise, fall and rise again, compares each with the CPU, and requires no new allocations once the pool has grown.
  • Full CUDA suite: passes on the RTX PRO 6000 (including the pooled W8A8 path) and on the RTX 3070 (24 packages each).
  • Pure-Go build: go build, go vet and go test ./... pass.

🤖 Generated with Claude Code

Cuts the host-side overhead out of long prefills. A 4314-token Nemotron 3 Nano prefill on the RTX PRO 6000 goes from **2.39 s to 1.40 s**, and the GPU is now busy 86% of the time, up from about 42%. ## What the profile showed After #91, the prefill's Nsight trace showed the GPU busy for only about 1.0 s of 2.39 s. Splitting its idle time by what the host was doing showed two costs, together about 1.2 s: | per prefill | idle time | cause | |---|---|---| | MoE routing | ~0.75 s | After each MoE layer's router logits came back to the host, the GPU waited a mean **32 ms** while every token was routed one at a time. Each token did a full stable sort of all 128 experts and six heap allocations: 5.4 µs per token. | | matmul scratch | ~0.47 s | Each large NVFP4 or bf16 matmul `cudaMalloc`'d and `cudaFree`'d a bf16 activation copy and an unpacked weight: about **10,600 pairs** per prefill, and `cudaFree` synchronizes. | | launch/sync, other host | ~0.15 s | about 12,000 ops, each synchronizing before the next launch (not addressed here) | ## Changes **Routing.** - `moe.SigmoidRouter.RouteInto` picks the top k directly: linear in the expert count, ties still broken toward the smaller index. It writes into caller-owned space and allocates nothing. `Route` wraps it, so GLM is unchanged. **0.79 µs per token**, down from 5.4 µs. - `nemotronh` routes a step's tokens across goroutines once there are 256 or more (a decode step stays on the calling goroutine). The packing then reads the picks back instead of routing inline. **Scratch pool (CUDA).** - The backend keeps a grow-only pool with one buffer per purpose: the bf16 activation, the unpacked bf16 or f32 weight, and W8A8's quantized activation and its scales. The matmul paths take views of these instead of allocating. Each buffer grows by at least half again when it grows, so rising sizes don't reallocate on every call. - Reuse is safe because every op finishes on the device before it returns. - The buffers still count as live allocations for `backend.MemStats`, as the transient scratch did before, and `Close` releases them. ## Results The same 4314-token prefill, measured with the same Nsight setup as #91: | | before | after | |---|---|---| | wall time | 2.39 s | **1.40 s** | | GPU idle | ~1.45 s | **0.19 s** | | `cudaMalloc`/`cudaFree` pairs | ~10,700 | 164 | | routing stall per MoE layer | ~32 ms | ~2.3 ms | | GPU busy share of wall time | ~42% | **86%** | - **Output:** unchanged, and chat still answers correctly. - **Now GPU-bound.** The remaining time is kernels: paged attention 0.47 s (the decode flash kernel on a 4k prompt), the NVFP4 unpack 0.22 s, the NVFP4 GEMV 0.17 s, scan and conv 0.14 s. - **Decode is unchanged at about 23 ms/token.** Its cost lies elsewhere, most likely each op's synchronization. It has not been profiled yet. ## Testing - **moe:** - `TestRouteIntoMatchesReference` checks `RouteInto` against a full-sort reference over 8,000 tokens. The logits are coarse so ties are common; the cases also include group masking that drives experts to −inf, and a k larger than the expert count. - `TestRouteIntoDoesNotAllocate` checks that routing a token allocates nothing. - **nemotronh:** `TestRouteParallelMatchesSerial` routes 1000 tokens through the parallel path and compares each with `Route`; it passes under `-race`. - **cuda:** `TestCUDAScratchReuse` runs NVFP4 matmuls whose sizes rise, fall and rise again, compares each with the CPU, and requires no new allocations once the pool has grown. - **Full CUDA suite:** passes on the RTX PRO 6000 (including the pooled W8A8 path) and on the RTX 3070 (24 packages each). - **Pure-Go build:** `go build`, `go vet` and `go test ./...` pass. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
perf: cut host overhead from long prefills (routing, matmul scratch)
All checks were successful
ci / test_and_build (pull_request) Successful in 27s
dfbb03929c
After #91, a profiled 4314-token Nemotron 3 Nano prefill took 2.39 s,
with the GPU busy for only about 1.0 s of it. Splitting the idle time
from the Nsight trace showed two host costs, together about 1.2 s:

- MoE routing: ~0.75 s. After each MoE layer's router logits came back
  to the host, the GPU sat idle a mean 32 ms while every token was
  routed, one at a time: a full stable sort of all 128 experts and six
  heap allocations per token, 5.4 us each.
- Matmul scratch: ~0.47 s. Each large NVFP4 or bf16 matmul cudaMalloc'd
  and cudaFree'd a bf16 activation copy and an unpacked weight, about
  10,600 pairs per prefill. cudaFree synchronizes.

Routing: moe.SigmoidRouter.RouteInto picks the top k directly (linear
in the expert count, ties still to the smaller index) into caller-owned
space, with no allocation; Route wraps it, so GLM is unchanged. It takes
0.79 us per token. nemotronh routes a step's tokens across goroutines
once there are 256 or more, and the packing reads the picks back
instead of routing inline.

Scratch: the CUDA backend keeps a grow-only pool with one buffer per
purpose (bf16 activation, unpacked bf16 or f32 weight, W8A8's quantized
activation and its scales), and the matmul paths take views of it
instead of allocating. Reuse is safe because every op finishes on the
device before returning. The buffers still count as live allocations
for backend.MemStats, as transient scratch did before, and Close
releases them.

Result on the same prefill: 2.39 s -> 1.40 s. GPU idle drops from
about 1.45 s to 0.19 s, with 164 cudaMalloc/cudaFree pairs where there
were about 10,700, and the routing stall per MoE layer drops from about
32 ms to 2.3 ms. Output is unchanged, and chat still answers correctly.
Decode is unchanged at about 23 ms/token; its cost lies elsewhere.

Tests:
- moe: TestRouteIntoMatchesReference checks RouteInto against a
  full-sort reference over 8,000 tokens, with coarse logits so ties are
  common, group masking that drives experts to -inf, and a k larger
  than the expert count. TestRouteIntoDoesNotAllocate checks it
  allocates nothing.
- nemotronh: TestRouteParallelMatchesSerial routes 1000 tokens through
  the parallel path and compares each with Route; it passes under
  -race.
- cuda: TestCUDAScratchReuse runs NVFP4 matmuls whose sizes rise and
  fall, compares each with the CPU, and requires no new allocations
  once the pool has grown.
- The CUDA suite passes on the RTX PRO 6000 (including the pooled W8A8
  path) and an RTX 3070.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator

Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id r-b9676a-e37583
This is an AI-generated review and may contain mistakes.

Status: ❌ Failed


This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id r-b9676a-e37583.

Comment @pr-reviewer-bot retry to try again.

<!-- pr-reviewer:review --> *Automated review by [pr-reviewer](https://git.brooktrails.org/brooktrails/pr-reviewer) v0.52.3 | Safety Check | Mistral Small | tracking id `r-b9676a-e37583`* *This is an AI-generated review and may contain mistakes.* **Status:** ❌ Failed --- This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id `r-b9676a-e37583`. Comment `@pr-reviewer-bot retry` to try again.
rcsheets deleted branch perf/host-overhead 2026-09-27 19:08:25 +00:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm!92
No description provided.