10 GiB of GPU memory unaccounted for after serving Nemotron 3 Nano #89
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observed
Serving NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 on the RTX PRO 6000 (94.5 GiB free at startup) with
--kv-cache 8GiB --max-model-len 4096, after two short chat requests had completed,/v1/internal/statusreported:weights_byteskv_cache_bytesstate_cache_bytesdevice_free_bytesWeights, KV and state account for 21.5 GiB, which would leave about 73 GiB free. The device reported about 10 GiB less than that, and nothing in gllm's accounting names it. The
gllm_memory_scratch_bytesgauge (allocated minus weights, KV and state) is the number that should explain the gap, but it was not captured during this run.(At the time the weights were widened to F32. Keeping BF16 resident, as of #88, brings the weights to 18.0 GiB, but that does not bear on the gap.)
Candidates
nemotronhforward allocates and frees its buffers per step throughscratch. A leak would grow with every step, so it should show up asscratch_bytesrising across requests.stage_i32/stage_f32cudaMalloc/cudaFree on every launch) are invisible toAllocBytes. None of them should be near 10 GiB, but the context and the lazily loaded kernel modules grow with the number of distinct kernels.nvfp4GEMVMaxRowsrows,matMulNVFP4dequantizes the weight to an F32 buffer for a cuBLASLt GEMM. The 28-token prefill of the first request may have taken that path for every expert, and a buffer that is cached instead of freed would be sized by the largest weight dequantized.To pin it down
--log-perfstats-interval 5sand comparescratch,device_usedandnvidia-smi's used memory:scratchstays near zero whiledevice_usedclimbs, the memory is outside gllm's allocator: look at the staging buffers and cuBLASLt first.This matters for sizing:
auto:fillsizes the KV cache from the free memory it sees after load. Memory that disappears after the first requests is memory the cache was promised, and on a tight fit it becomes an OOM mid-serving.