Greedy outputs depend on batch composition (NVFP4 GEMV/GEMM switch at 32 rows is a candidate) #98
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observation
At temperature 0, gllm returns different outputs for the same prompt depending on what else is in the batch. Requests served one at a time are reproducible.
Measured on NVIDIA-Nemotron-3-Nano-30B-A3B-NVFP4 (gllm
51ef3b7, RTX PRO 6000 Blackwell) withcmd/glbench/configs/nemotron3-vs-vllm.yaml. Each workload cycles 8 distinct prompts, so every prompt is sent several times per run.Comparing the two runs request by request, 8 of 32 outputs differ at concurrency 8 and 37 of 64 at concurrency 32. The outputs stay coherent and diverge only partway through, when a token near a tie flips. One example, "Give three tips for debugging a CUDA out-of-memory error.", diverged at character 248:
For reference, vLLM 0.30.0 (default settings) in the same run also varies across batches: all 8 prompts produced more than one output, with up to 8 variants of a single prompt. Some batch dependence is normal for batched inference, so this is probably not a correctness bug. It is worth deciding on explicitly, though, because it:
Candidate cause (unverified)
matMulNVFP4(internal/backend/cuda/matmul.go) picks its kernel by row count:n <= nvfp4GEMVMaxRows(32): fused-dequant GEMV with f32 activationsIn the gathered MoE path,
nfor an expert's matmul is the number of (token, expert) pairs routed to that expert in the step. That depends on which other sequences share the step, so the same token's expert FFN can run at f32 or bf16 activation precision depending on its neighbours. Decode steps at concurrency 8 stay below 32 rows per expert. A prefill that shares a step with other prompts can cross the threshold, so the divergence would enter through prefill.Other things that may contribute: reduction order in any kernel whose work split depends on batch shape (split-K, the attention dispatch threshold), and the order in which gathered expert outputs are summed.
Next steps
🤖 Generated with Claude Code