10 GiB of GPU memory unaccounted for after serving Nemotron 3 Nano #89

Open
opened 2026-09-27 13:02:48 +00:00 by rcsheets · 0 comments
Owner

Observed

Serving NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 on the RTX PRO 6000 (94.5 GiB free at startup) with --kv-cache 8GiB --max-model-len 4096, after two short chat requests had completed, /v1/internal/status reported:

GiB
weights_bytes 20.02
kv_cache_bytes 0.75
state_cache_bytes 0.74
device_free_bytes 63.26

Weights, KV and state account for 21.5 GiB, which would leave about 73 GiB free. The device reported about 10 GiB less than that, and nothing in gllm's accounting names it. The gllm_memory_scratch_bytes gauge (allocated minus weights, KV and state) is the number that should explain the gap, but it was not captured during this run.

(At the time the weights were widened to F32. Keeping BF16 resident, as of #88, brings the weights to 18.0 GiB, but that does not bear on the gap.)

Candidates

  • Forward scratch that is never freed. The nemotronh forward allocates and frees its buffers per step through scratch. A leak would grow with every step, so it should show up as scratch_bytes rising across requests.
  • CUDA-side caching outside gllm's allocator. The cuBLASLt workspace (32 MiB), the CUDA context, and per-launch staging buffers (stage_i32/stage_f32 cudaMalloc/cudaFree on every launch) are invisible to AllocBytes. None of them should be near 10 GiB, but the context and the lazily loaded kernel modules grow with the number of distinct kernels.
  • The NVFP4 dequant scratch. Above nvfp4GEMVMaxRows rows, matMulNVFP4 dequantizes the weight to an F32 buffer for a cuBLASLt GEMM. The 28-token prefill of the first request may have taken that path for every expert, and a buffer that is cached instead of freed would be sized by the largest weight dequantized.

To pin it down

  1. Serve the same checkpoint with --log-perfstats-interval 5s and compare scratch, device_used and nvidia-smi's used memory:
    • right after load
    • after a first prefill
    • after several decode steps
    • after many requests
  2. If scratch stays near zero while device_used climbs, the memory is outside gllm's allocator: look at the staging buffers and cuBLASLt first.
  3. If it steps up once and then holds, it is a cache, sized by the first large prefill. If it climbs per request, it is a leak.
  4. Check whether other architectures (a Mistral checkpoint of similar size) show the same gap. That separates the new hybrid code paths from the shared CUDA backend.

This matters for sizing: auto:fill sizes the KV cache from the free memory it sees after load. Memory that disappears after the first requests is memory the cache was promised, and on a tight fit it becomes an OOM mid-serving.

## Observed Serving NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 on the RTX PRO 6000 (94.5 GiB free at startup) with `--kv-cache 8GiB --max-model-len 4096`, after two short chat requests had completed, `/v1/internal/status` reported: | | GiB | |---|---| | `weights_bytes` | 20.02 | | `kv_cache_bytes` | 0.75 | | `state_cache_bytes` | 0.74 | | `device_free_bytes` | 63.26 | Weights, KV and state account for 21.5 GiB, which would leave about 73 GiB free. The device reported about 10 GiB less than that, and nothing in gllm's accounting names it. The `gllm_memory_scratch_bytes` gauge (allocated minus weights, KV and state) is the number that should explain the gap, but it was not captured during this run. (At the time the weights were widened to F32. Keeping BF16 resident, as of #88, brings the weights to 18.0 GiB, but that does not bear on the gap.) ## Candidates - **Forward scratch that is never freed.** The `nemotronh` forward allocates and frees its buffers per step through `scratch`. A leak would grow with every step, so it should show up as `scratch_bytes` rising across requests. - **CUDA-side caching outside gllm's allocator.** The cuBLASLt workspace (32 MiB), the CUDA context, and per-launch staging buffers (`stage_i32`/`stage_f32` cudaMalloc/cudaFree on every launch) are invisible to `AllocBytes`. None of them should be near 10 GiB, but the context and the lazily loaded kernel modules grow with the number of distinct kernels. - **The NVFP4 dequant scratch.** Above `nvfp4GEMVMaxRows` rows, `matMulNVFP4` dequantizes the weight to an F32 buffer for a cuBLASLt GEMM. The 28-token prefill of the first request may have taken that path for every expert, and a buffer that is cached instead of freed would be sized by the largest weight dequantized. ## To pin it down 1. Serve the same checkpoint with `--log-perfstats-interval 5s` and compare `scratch`, `device_used` and `nvidia-smi`'s used memory: - right after load - after a first prefill - after several decode steps - after many requests 2. If `scratch` stays near zero while `device_used` climbs, the memory is outside gllm's allocator: look at the staging buffers and cuBLASLt first. 3. If it steps up once and then holds, it is a cache, sized by the first large prefill. If it climbs per request, it is a leak. 4. Check whether other architectures (a Mistral checkpoint of similar size) show the same gap. That separates the new hybrid code paths from the shared CUDA backend. This matters for sizing: `auto:fill` sizes the KV cache from the free memory it sees after load. Memory that disappears after the first requests is memory the cache was promised, and on a tight fit it becomes an OOM mid-serving.
Sign in to join this conversation.
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
brooktrails/gllm#89
No description provided.