perf: cut host overhead from long prefills (routing, matmul scratch) #92
Loading…
Reference in a new issue
No description provided.
Delete branch "perf/host-overhead"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Cuts the host-side overhead out of long prefills. A 4314-token Nemotron 3 Nano prefill on the RTX PRO 6000 goes from 2.39 s to 1.40 s, and the GPU is now busy 86% of the time, up from about 42%.
What the profile showed
After #91, the prefill's Nsight trace showed the GPU busy for only about 1.0 s of 2.39 s. Splitting its idle time by what the host was doing showed two costs, together about 1.2 s:
cudaMalloc'd andcudaFree'd a bf16 activation copy and an unpacked weight: about 10,600 pairs per prefill, andcudaFreesynchronizes.Changes
Routing.
moe.SigmoidRouter.RouteIntopicks the top k directly: linear in the expert count, ties still broken toward the smaller index. It writes into caller-owned space and allocates nothing.Routewraps it, so GLM is unchanged. 0.79 µs per token, down from 5.4 µs.nemotronhroutes a step's tokens across goroutines once there are 256 or more (a decode step stays on the calling goroutine). The packing then reads the picks back instead of routing inline.Scratch pool (CUDA).
backend.MemStats, as the transient scratch did before, andClosereleases them.Results
The same 4314-token prefill, measured with the same Nsight setup as #91:
cudaMalloc/cudaFreepairsTesting
TestRouteIntoMatchesReferencechecksRouteIntoagainst a full-sort reference over 8,000 tokens. The logits are coarse so ties are common; the cases also include group masking that drives experts to −inf, and a k larger than the expert count.TestRouteIntoDoesNotAllocatechecks that routing a token allocates nothing.TestRouteParallelMatchesSerialroutes 1000 tokens through the parallel path and compares each withRoute; it passes under-race.TestCUDAScratchReuseruns NVFP4 matmuls whose sizes rise, fall and rise again, compares each with the CPU, and requires no new allocations once the pool has grown.go build,go vetandgo test ./...pass.🤖 Generated with Claude Code
Automated review by pr-reviewer v0.52.3 | Safety Check | Mistral Small | tracking id
r-b9676a-e37583This is an AI-generated review and may contain mistakes.
Status: ❌ Failed
This review couldn't be completed: that model isn't loaded on the inference service right now, and no alternate model was able to review it either. Consider splitting this PR into smaller changes. Tracking id
r-b9676a-e37583.Comment
@pr-reviewer-bot retryto try again.